Safeguards Enforcement Analyst, User Well-being

$245k - $285k Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC

Posted 1d ago

Job Location

Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC

Tech Stack

Remote Work Policy

On-site

Categories

Applied AI Engineer

About the job

As a Safeguards Analyst on the User Well-being team, you will support the design and deployment of mental health guardrails. This involves iterating on detection systems, managing review queues, evaluating new interventions, and monitoring existing ones. The work translates expert clinical guidance, data analyses, and constraints into concrete changes for detection, review, and response. The team addresses harms including suicide, self-harm, disordered eating, AI sycophancy, and emotional dependence on AI, with potential to expand into broader user well-being enforcement. Safety is central to the mission, and this role will help shape policy enforcement for harmless, helpful, and honest user interactions with AI products.

Responsibilities

  • Support the design and execution of interventions, defining key metrics, and curating evaluation datasets.
  • Partner with Engineering and Data Science teams to build, tune, and validate detection models for automated intervention systems.
  • Monitor the performance of interventions and detection systems over time.
  • Review flagged content to drive enforcement and policy improvements.
  • Support the development of in-product features that connect users to crisis resources, working with Product, Legal, and external partners.
  • Provide detailed feedback on policy gaps to the Safeguards Policy Design team based on real scenarios.
  • Stay updated on emerging AI policy and research related to AI and mental health to inform decision-making.
  • Translate policy definitions into measurable forms like rubrics, review guidelines, or classification criteria.

Requirements

  • Experience in trust & safety, product policy, content moderation, or a related field, with direct exposure to mental health or related well-being harm areas.
  • Experience designing or running experiments, evaluations, or measurement studies.
  • Proficiency in SQL and/or other data analysis tools.
  • Experience working with generative AI products, including writing effective prompts for content review, classification, or evaluation.
  • Experience turning open questions and data into concise and insightful analysis.
  • Experience identifying emerging risks and communicating findings to cross-functional stakeholders.
  • Understanding of the challenges in implementing product policies at scale in content moderation.
  • Sound judgment in ambiguous, high-consequence cases and comfort escalating appropriately.

Benefits

  • Annual compensation range: $245,000 - $285,000 USD

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.