Reinforcement Learning Jobs

51 open roles mentioning Reinforcement Learning

Research Engineer, RL Engineering

1d ago
Anthropic

Anthropic

Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. As an ML Systems Engineer on the Reinforcement Learning Engineering team, you will build and improve the critical algorithms and infrastructure that researchers use to train AI models like Claude. Your work will directly enable breakthroughs in AI capabilities and safety, focusing on enhancing the performance, robustness, and usability of these systems to accelerate research progress. You will support and empower the research team in their mission to build beneficial AI systems, specifically by building, maintaining, and improving the algorithms and systems used for finetuning production and research models with methods like RLHF.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicPythonFine-Tuning +3 more

[Expression of Interest] Research Engineer / Scientist, Alignment - London

2d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to ensure AI is safe and beneficial for society. As a Research Engineer on the Alignment Science team in London, you will design and execute machine learning experiments to understand and steer the behavior of advanced AI systems. You will focus on AI safety, particularly risks from future human-level AI systems, collaborating with teams like Interpretability and Frontier Red Team. The role involves exploratory research in areas such as AI Control and Alignment Stress-testing, aiming to make AI helpful, honest, and harmless.

London, UK onsite
AnthropicKubernetesPython +3 more

Anthropic Fellows Program, Reinforcement Learning

2d ago
Anthropic

Anthropic

Anthropic is seeking talented individuals for its Fellows Program, focusing on AI research and engineering. This program provides funding and mentorship to promising technical talent, regardless of prior experience, to work on empirical projects aligned with Anthropic's research priorities. The goal is to produce public outputs, such as research papers, contributing to the development of reliable, interpretable, and steerable AI systems that are safe and beneficial for society. Fellows will engage in full-time research for four months, with opportunities for extension, and will be mentored by Anthropic researchers.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +3 more

Anthropic Fellows Program, ML Systems & Performance

2d ago
Anthropic

Anthropic

Anthropic is seeking talented individuals for its Fellows Program, focusing on AI research and engineering to ensure AI systems are reliable, interpretable, and steerable. This program provides funding and mentorship to promising technical talent, regardless of prior experience, to work on empirical projects aligned with Anthropic's research priorities. The goal is to produce public outputs, such as paper submissions, contributing to the development of safe and beneficial AI for society. Fellows will engage in full-time research for four months, with direct mentorship from Anthropic researchers and access to a shared workspace in either Berkeley, California, or London, UK. The program also offers connections to the broader AI safety and security research community.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +3 more

Anthropic Fellows Program, AI Security

2d ago
Anthropic

Anthropic

Anthropic is seeking talented individuals for the Anthropic Fellows Program, focusing on AI Security. This program is designed to foster AI research and engineering talent by providing funding and mentorship to promising individuals, regardless of prior experience. Fellows will work on empirical projects using external infrastructure, aiming to produce public outputs like research papers. The program offers a structured 4-month full-time research period with direct mentorship from Anthropic researchers, access to a shared workspace in Berkeley or London, and connections to the AI safety and security research community.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +3 more

Anthropic Fellows Program, AI Safety

2d ago
Anthropic

Anthropic

Anthropic is seeking talented individuals for its Fellows Program, focusing on AI Safety. This program aims to foster AI research and engineering talent by providing funding and mentorship to promising individuals, regardless of prior experience. Fellows will engage in empirical projects using external infrastructure and open-source models, with the goal of producing public outputs like research papers. The program offers a structured 4-month full-time research period, direct mentorship from Anthropic researchers, access to shared workspaces in Berkeley or London, and connections to the AI safety community. The program is designed to encourage diverse perspectives and applications, even from those who may not meet every single qualification.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +1 more

Anthropic Fellows Program

2d ago
Anthropic

Anthropic

Anthropic is seeking candidates for its Fellows Program, designed to cultivate AI research and engineering talent. This program provides funding and mentorship to promising individuals, irrespective of prior experience, to work on empirical projects aligned with Anthropic's research priorities. The goal is to produce public outputs, such as paper submissions, with a strong track record of fellows achieving this in previous cohorts. The program offers a 4-month full-time research period, direct mentorship from Anthropic researchers, access to a shared workspace in Berkeley or London, and connections to the AI safety and security research community. Applications are reviewed on a rolling basis for cohorts starting in July 2026 and beyond.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +5 more

Staff Research Engineer, Discovery Team

2d ago
Anthropic

Anthropic

Anthropic is dedicated to building reliable, interpretable, and steerable AI systems that are safe and beneficial for society. As a Research Engineer on the Discovery Team, you will work end-to-end to identify and address key blockers on the path to scientific Artificial General Intelligence (AGI). This role involves improving models' abilities to use computers, acting as a laboratory for long-horizon tasks and a crucial component for scientific workflows. You will collaborate with a team of researchers and engineers focused on pushing the scientific frontier.

San Francisco, CA onsite
AnthropicDockerKubernetes +2 more

Research Engineer, Universes

2d ago
Anthropic

Anthropic

The Universes team within Research is responsible for training AI models to perform complex, difficult, long-horizon agentic tasks in ultra-realistic settings. We design and implement novel training environments that go far beyond what models can do today — environments where models learn to navigate ambiguity, handle interruptions, maintain context over extended interactions, and exercise judgment in open-ended scenarios. We're looking for Research Engineers to help us build the next generation of training environments for capable and safe agentic AI. This role blends research and engineering responsibilities, requiring you to both implement novel approaches and contribute to research direction. You'll work on fundamental research in reinforcement learning, designing training environments and methodologies that push the state of the art, and building evaluations that measure genuine capability.

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicGoFine-Tuning +1 more

ML/Research Engineer, Safeguards

2d ago
Anthropic

Anthropic

Anthropic is seeking ML Engineers and Research Engineers to join the Safeguards ML team. The primary focus of this role is to develop systems that detect and mitigate misuse of AI systems, ranging from individual policy violations to sophisticated coordinated attacks. You will build defenses to ensure product safety as AI capabilities advance, protect user well-being, and guarantee appropriate model behavior across various contexts. This work is crucial for Anthropic's Responsible Scaling Policy commitments.

San Francisco, CA | New York City, NY onsite
AnthropicPythonReinforcement Learning +1 more

Research Engineer / Scientist, Alignment

2d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer/Scientist for its Alignment Science team. This role involves designing and executing machine learning experiments to understand and steer the behavior of advanced AI systems, with a focus on AI safety and potential risks from future human-level AI. You will collaborate with other teams on exploratory research, contributing to Anthropic's mission of creating reliable, interpretable, and steerable AI systems that are helpful, honest, and harmless.

San Francisco, CA onsite
AnthropicKubernetesPython +4 more

Research Engineer/Research Scientist, Pre-training

2d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer to join its Pre-training team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, contributing to the creation of safe, steerable, and trustworthy AI systems. The team is dedicated to ensuring that transformative AI systems are aligned with human interests and societal benefit.

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicKubernetesPython +4 more

Research Engineer, Pretraining

2d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer to join its Pretraining team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, aiming to build safe, steerable, and trustworthy AI systems. The mission is to ensure that transformative AI systems are aligned with human interests and are beneficial for society. The team is dedicated to pushing the boundaries of AI while prioritizing safety and ethics.

London, UK onsite
AnthropicKubernetesPython +4 more

Research Engineer, Performance RL (Reinforcement Learning)

2d ago
Anthropic

Anthropic

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the RL Teams Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.6 and Opus 4.6. Our work spans several key areas: - Developing systems that enable models to use computers effectively - Advancing code generation through reinforcement learning - Pioneering fundamental RL research for large language models - Building scalable RL infrastructure and training methodologies - Enhancing model reasoning capabilities We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish. About the Role We're hiring for the Code RL team within the RL organization. As a Research Engineer, you'll advance our models' ability to safely write correct, fast code for accelerators. You'll need to know accelerator performance well to turn it into tasks and signals models can learn from. Specifically, you will: - Invent, design and implement RL environments and evaluations. - Conduct experiments and shape our research roadmap. - Deliver your work into training runs. - Collaborate with other researchers, engineers, and performance engineering specialists across and outside Anthropic. You may be a good fit if you: - Have expertise with accelerators (CUDA, ROCm, Triton, Pallas), ML framework programming (JAX or PyTorch). - Have worked across the stack – kernels, model code, distributed systems. - Know how to balance research exploration with engineering implementation. - Are passionate about AI's potential and committed to developing safe and beneficial systems. Strong candidates may also have: - Experience with reinforcement learning. - Experience porting ML workloads between different types of accelerators. - Familiarity with LLM training methodologies. The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $350,000—$850,000 USD Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings. How we're different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come work with us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.

San Francisco, CA onsite
AnthropicPyTorchClaude +2 more

Research Engineer, Knowledge Team

2d ago
Anthropic

Anthropic

Anthropic is seeking Research Engineers to reimagine how Claude interacts with external data sources. This role involves designing novel architectures for organizing information and training language models to effectively utilize these architectures. The goal is to move beyond traditional data paradigms to accommodate the capabilities of Large Language Models (LLMs).

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicPythonRAG +3 more

Research Engineer, Cybersecurity RL (Reinforcement Learning)

2d ago
Anthropic

Anthropic

The Cybersecurity RL team within Anthropic is seeking a Research Engineer to advance the capabilities of AI models in secure coding, vulnerability remediation, and defensive cybersecurity. This role combines research and engineering, requiring the design and implementation of RL environments, conducting experiments, delivering work into production training runs, and collaborating with cross-functional teams. The ideal candidate will have domain expertise in cybersecurity and an interest or experience in training safe AI models, potentially coming from backgrounds like white hat hacking, security engineering, or detection and response.

San Francisco, CA | New York City, NY onsite
AnthropicClaudeReinforcement Learning

Research Engineer, Code RL (Reinforcement Learning)

2d ago
Anthropic

Anthropic

We are seeking a Research Engineer for our Code RL team, focused on advancing AI models' capabilities in writing, editing, testing, debugging, and shipping real software. This role involves designing RL environments, coding tasks, and reward signals, as well as running training experiments on frontier models. You will diagnose model performance, improve pipeline speed and reliability, and contribute to areas like agentic coding behaviors, code correctness, and autonomous engineering. The position blends cutting-edge research with practical engineering to build high-quality, scalable AI systems.

San Francisco, CA | New York City, NY onsite
AnthropicPythonFine-Tuning +5 more

Model Performance Software Engineer, Claude Code

2d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. We are a growing team of researchers, engineers, policy experts, and business leaders dedicated to this mission. We are seeking a Staff Software Engineer to lead technical direction at the intersection of engineering and research for the Claude Code team. In this role, you will collaborate with researchers and engineering leadership to define how we measure, understand, and enhance Claude's coding abilities. You will architect the systems, tooling, and evaluation infrastructure that accelerate our research progress and be responsible for technical decisions impacting the team and beyond. This senior individual contributor position is for someone with a proven track record of building and owning large-scale systems, ready to take on a technical leadership role by driving architecture, mentoring engineers, and influencing the future of Claude Code.

San Francisco, CA | New York City, NY onsite
AnthropicPythonTypeScript +2 more

Research Engineer, Domain Scaling

2d ago
Anthropic

Anthropic

The Domain Scaling team aims to make Claude world-class at real-world knowledge work in domains like finance, healthcare, and legal. This role combines direct applied research with data sourcing (real-world and synthetic) to improve our models. You will own the end-to-end process of creating RL environments for new capabilities, which includes identifying high-value tasks, designing reward signals, managing vendor relationships, and measuring impact on model performance.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicFine-TuningClaude +1 more

Software Engineer, RL Data

2d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. This senior, foundational role on a new team involves making key architectural decisions and shaping initial development. The work is hands-on and varied, encompassing pipeline and infrastructure engineering, prompt tuning, and supporting research teams. The RL Data team focuses on building systems for high-quality reinforcement learning data for Claude, including data collection pipelines, human feedback tooling, execution environments, and quality assurance to ensure trustworthy training data at scale. The goal is to enhance Claude's capabilities in real-world tasks, particularly in AI safety research and beneficial AI deployments.

San Francisco, CA | New York City, NY onsite
AnthropicDockerKubernetes +4 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.