Head of Policy Design, Societal Harms
- Location
- San Francisco, CA
- Type
- full time
- Posted
- Aug 28, 2026
<div class="content-intro"><h2><strong>About Anthropic</strong></h2> <p>Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.</p></div><h2><strong>About the role</strong></h2> <p>Anthropic's Safeguards organization builds the policies, evaluations, and detection and enforcement systems that define and hold the limits on how Claude can be used. In this role, you'll lead our policy design team, managing the teams responsible for radicalization, child safety, user well-being, harmful manipulation, and election integrity, among other harm areas.</p> <p>The team is responsible for understanding and defining the risks that come with engaging with Claude, how those risks materialize in the real world, and the mitigations needed to prevent them. As the manager, you'll work with your team to draw the boundaries between what is and is not allowed, then partner with research, product, and engineering to build the right interventions. Mitigating these harms takes the whole stack: the values and judgment trained into the model itself, the policies and detection systems we enforce on top of it, and the interventions we build into our products. More capable models, new product surfaces, and new user behaviors will keep testing these boundaries, so the team's policies have to keep pace.</p>…
Looking for more like this? Browse all AI Agent Engineer jobs.