Reinforcement Learning Jobs
51 open roles mentioning Reinforcement Learning
Anthropic Fellows Program, The Anthropic Institute (Economics & Policy)
Anthropic
Anthropic is seeking talented individuals for its Fellows Program, designed to foster AI research and engineering talent. The program provides funding and mentorship to promising technical individuals, regardless of prior experience, to work on empirical projects aligned with Anthropic's research priorities. The goal is to produce public outputs, such as paper submissions. Fellows will primarily utilize external infrastructure like open-source models and public APIs. The program emphasizes AI safety and beneficial AI development for society.
$25k - $50k
Research Engineer, Life Sciences
Anthropic
Anthropic is seeking an exceptional Research Engineer to join its Life Sciences team. This role focuses on accelerating progress in life sciences through AI, from early discovery to translation. You will leverage deep expertise in machine learning engineering to develop novel evaluation frameworks and training strategies, pushing the boundaries of AI in biology. Working at the intersection of AI and biological sciences, you will develop rigorous methods to measure and improve model performance on complex scientific tasks, collaborating with researchers and engineers to build AI systems for all phases of research and development, while upholding Anthropic's commitment to safety and beneficial impact. Previous experience in life sciences is welcome but not required.
Research Engineer, Computer Use
Anthropic
The Computer Use team focuses on teaching Claude to see, use, and understand computer interfaces. As a Research Engineer on the team, you'll work on advancing our models' ability to reliably and safely operate real software. We're looking for someone who's genuinely excited about both the research and the product sides of computer use. Your work will translate directly into model improvements in our own and our customers' products. You can try Claude's computer use capabilities today through the Claude in Chrome extension and Claude Cowork.
Research Engineer, Visual Knowledge Work
Anthropic
We are seeking research engineers with a strong computer vision background to enhance the visual and spatial reasoning capabilities of our state-of-the-art Claude models. This role involves research, development, and evaluation, taking a full-stack approach across pretraining, RL, and runtime techniques. You will collaborate closely with the product organization to ensure that vision improvements directly impact Claude's performance on real-world tasks and address customer challenges.
Research Engineer, Chip Design RL (Reinforcement Learning)
Anthropic
Anthropic is seeking a Research Engineer for its Code RL team to advance AI models' ability to design silicon. This role involves leveraging chip design expertise to create tasks and signals for models, focusing on hardware design domains. The position is at the intersection of cutting-edge research and engineering excellence, with a commitment to building high-quality, scalable systems that push the boundaries of AI capabilities.
Staff Software Engineer, Code RL
Anthropic
Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. This role focuses on the engineering aspects of reinforcement learning for Claude's coding capabilities, involving the creation and scaling of agentic coding environments. You will have significant influence on technical direction and standards, embedding with research teams to understand their needs, build supporting frameworks and infrastructure, and then transfer ownership of these well-maintained systems. The role also includes ensuring the ongoing health and maintainability of production RL runs, including monitoring and triage tooling. The team's work spans client-side sandboxed execution for agentic RL, large-scale data processing, dataset lifecycle management, and the frameworks researchers use to build environments. You will focus on areas where your deep expertise is most valuable, particularly if you have strong Python skills, a keen eye for API and framework design, and experience with complex system failures.
Staff Software Engineer, Environments Infrastructure
Anthropic
Anthropic's Environments organization is responsible for building and maintaining the infrastructure that enhances Claude's capabilities through reinforcement learning. This includes developing frameworks for researchers to create environments and managing the infrastructure that runs them, with a core mission to productionize research. The role involves embedding with research teams to understand their workflows and designing frameworks and APIs that accelerate their progress, ensuring these systems are understandable, ownable, and maintainable. A key aspect is also ensuring the health, maintainability, monitoring, and ease of triage for production RL runs.
Research Engineer, Developer Experience, Tinker
thinkingmachines
Thinking Machines Lab is seeking a Research Engineer focused on developer experience to build and enhance their Tinker platform. This role involves working hands-on with users to understand their challenges and translate them into product improvements. You will be responsible for creating and updating documentation, adding library features, prototyping integrations, and ensuring users can smoothly customize frontier AI models. This position acts as a crucial link between Tinker users and the internal research and infrastructure teams, surfacing user patterns to inform product and infrastructure priorities and sharing learnings through various channels.
$350k - $475k
Research Lead, Training Insights
Anthropic
As a Research Lead on the Training Insights team, you will develop the strategy for, and lead execution on, how we measure and characterize model capabilities across training and deployment. This is a hands-on leadership role where you will drive original research into new evaluation methodologies while leading a small team of researchers and research engineers. Your work will span the full lifecycle of model development, from researching and building new long-horizon evaluations to developing novel approaches for measuring emerging capabilities and deepening our understanding of how those capabilities develop. You will also take a cross-organizational view, working across various teams to map the landscape of model evaluations and identify critical gaps. This role carries significant visibility and impact, helping to shape the evaluation narrative for model releases and contributing directly to how Anthropic communicates about its models. Done well, you will change how the industry measures and understands model capabilities, significantly furthering our safety mission.
Member of Technical Staff - Post-Training
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible for everyone to use, customize, and build on. We are building open models that empower individuals to control their intelligence and shape the future of AI. As a Member of Technical Staff - Post-Training, you will play a crucial role in transforming powerful pre-trained models into aligned and general agents. This position involves driving research and engineering initiatives at the forefront of post-training techniques, from data curation to large-scale optimization, and contributing to the advancement of large model reasoning and instruction following capabilities.
Senior AI Product Manager, Healthcare Agents
Scale AI
Scale is seeking an AI Product Manager to lead the Healthcare vertical within the Agents Data & Reinforcement Learning Environments team. This role involves owning the development of realistic RL environments for training and evaluating AI agents in healthcare software and workflows, as well as defining the "data as a product" strategy that supports them. The ideal candidate will possess deep understanding of the Healthcare industry and its workflows, combined with insight into AI research and current agent capabilities. You will translate this expertise into environments and datasets that enable AI agents to perform real healthcare tasks, serving as the domain expert for Scale's key customers and researchers. A strong entrepreneurial and go-to-market mindset is essential for success in this position.
$206k - $257k
Senior AI Product Manager, Finance Agents
Scale AI
Scale is seeking an AI Product Manager to lead the Finance vertical within the Agents Data & Reinforcement Learning Environments team. This role involves owning the development of realistic RL environments for training and evaluating AI agents in financial workflows, as well as defining the "data as a product" strategy that supports them. You will leverage your deep understanding of the Finance industry and AI capabilities to identify valuable financial tasks for AI modeling, determine data sourcing and structuring strategies, and translate domain expertise into a defensible product. The ideal candidate will possess a strong entrepreneurial and go-to-market mindset, coupled with the ability to pair finance industry experience with an understanding of AI research and current agent capabilities in financial workflows.
$205k - $257k
Research Intern RL & Post-Training Systems, Turbo (Fall 2026)
Together AI
The Turbo Research team focuses on making post-training and reinforcement learning for large language models efficient, scalable, and reliable. This work intersects RL algorithms, inference systems, and large-scale experimentation, where inference costs significantly impact training efficiency and the practicality of learning algorithms. As a research intern, you will investigate RL and post-training methods whose performance and scalability are closely tied to inference behavior, co-designing algorithms and systems. Projects aim to enable new experimental regimes, including larger models, longer rollouts, and more complex evaluations, by re-evaluating the interaction between inference, scheduling, and training.
Software Engineer, Identity
Scale AI
Scale is seeking a Software Engineer, Identity to join our Platform Engineering team. In this role, you will be instrumental in designing and developing core platforms and software systems, with a specific focus on identity, access management, authorization, and authentication. You will gain broad exposure to the cutting edge of the AI industry as Scale supports enterprises, startups, and governments. This position offers the opportunity to contribute to the foundational elements of products that power advanced LLMs and generative models, playing a crucial role in how humanity interacts with AI.
$216k - $270k
AI research scientist
Writer
AI research at WRITER focuses on building the scientific foundation for ambitious enterprise AI deployments. As a staff AI research scientist, you will drive a high-impact research agenda centered on large language models, agentic reasoning, and system-level capabilities essential for enterprise-scale AI. This role offers a unique opportunity to advance the field while directly contributing to products used by hundreds of thousands daily. You will work on post-training, planning, multi-step reasoning, and agentic workflows, directly shaping the future of enterprise AI performance and scalability. The role provides resources, infrastructure, and cross-functional support to pursue and implement ambitious ideas rapidly.
Research Intern, Model Shaping (Fall 2026)
Together AI
As a Research Intern in the Model Shaping team, you will work on advanced post-training methods, new techniques for efficient neural network training, and robust evaluation of foundation model capabilities. The Model Shaping team at Together AI focuses on tailoring open foundation models for downstream applications, building services for machine learning developers, and developing new methods for efficient model training and evaluation. This role offers the opportunity to contribute to cutting-edge research and potentially influence open-source projects.
Technical Program Manager, Engineering
Scale AI
Scale is at the forefront of the AI revolution, developing data engines and technologies that power the world's leading LLMs. This role focuses on leading critical programs within the Platform and Security Engineering teams, overseeing the design and development of core data storage systems and security initiatives. You will drive company-wide programs, improve processes, and ensure alignment with industry standards, gaining exposure to the cutting edge of AI adoption across various sectors. The work is crucial for making AI models safe, aligned, and useful through human evaluation and reinforcement learning.
$181k - $226k
Senior/Staff Machine Learning Research Engineer, General Agents, Enterprise GenAI
Scale AI
Scale AI is seeking a Senior/Staff Machine Learning Engineer for its General Agents team. This role is crucial in designing, building, and deploying production-ready AI agents to address high-impact enterprise challenges. You will be involved in the entire agent lifecycle, from conceptualization and system design to evaluation, deployment, and ongoing iteration. The position requires bridging cutting-edge agentic techniques with the practical demands of real-world customer environments, focusing on creating scalable, reliable, and generalizable agent systems.
$265k - $331k
Software Engineer, Platform
Scale AI
Scale is at the forefront of the AI revolution, building the Generative AI Data Engine and other products that power the world's most advanced LLMs. The Platform Engineering team is foundational to these efforts, responsible for designing and developing shared platforms, architecting core cloud infrastructure, and redefining software development processes. This role offers exposure to the cutting edge of AI development across various sectors, from startups to governments. You will drive the design and implementation of critical platforms, collaborate with cross-functional teams, and proactively improve engineering practices. This is an opportunity to shape the future of AI infrastructure and contribute to some of the most important work in how humanity interacts with AI.
$216k - $270k
Technical Lead Manager, Physical AI
Scale AI
Scale AI is seeking a Technical Lead Manager for its Physical AI team, focusing on the development of general AI that can reason and act in the physical world. This role bridges cutting-edge Machine Learning research with physical robot deployment, leading a team of Research Engineers while remaining a hands-on technical contributor. The primary focus is on developing and evaluating Large-Scale Foundation Models, such as VLAs and World models, to enable robots and autonomous vehicles to generalize across diverse tasks and environments. The team leverages Scale's extensive data infrastructure to help build Foundation Models for Physical AI, aiming to redefine the future of automation.
$249k - $311k