The Safeguards Enforcement Analyst will ensure AI models meet safety and policy standards through evaluations and mitigations, collaborating with cross-functional teams.
A weekly digest of AI safety, governance, policy and responsible AI roles.
By subscribing, you agree to receive the AI Safety Careers newsletter. We use MailerLite to send emails and may track opens and clicks to improve the newsletter. You can unsubscribe at any time. See our Privacy Policy.
677 active roles found
The Safeguards Enforcement Analyst will ensure AI models meet safety and policy standards through evaluations and mitigations, collaborating with cross-functional teams.
The Research Scientist, Interpretability role at Anthropic focuses on mechanistic interpretability to enhance the safety and understanding of AI systems.
The Research Engineer, Universes role at Anthropic focuses on developing training environments for safe AI systems and includes responsibilities for building evaluations to measure AI capabilities.
Get the best AI safety, governance, policy and responsible AI roles in your inbox.
By subscribing, you agree to receive the AI Safety Careers newsletter. We use MailerLite to send emails and may track opens and clicks to improve the newsletter. You can unsubscribe at any time. See our Privacy Policy.
The role involves conducting research on AI safety and alignment, focusing on understanding and steering the behavior of powerful AI systems.
Research Engineer/Scientist on Anthropic’s Alignment Science team, conducting experimental AI safety research on powerful future systems, safety evaluations, alignment stress-testing, and related safeguards work.
The role focuses on building reliable and interpretable AI systems, emphasizing safety and societal impacts.
The Research Engineer/Research Scientist role at Anthropic focuses on developing large language models with an emphasis on safety, alignment, and societal impacts.
The Research Engineer will enhance AI model safety and alignment through post-training techniques, impacting the quality and capabilities of production models.
The Research Engineer will focus on post-training processes to enhance AI model safety and alignment, implementing techniques to improve production model quality.
The Research Engineer will work on training and optimizing large-scale AI models, ensuring their reliability and safety, and addressing production issues.
The Research Engineer will work on training and optimizing production pretrained models, ensuring their reliability and efficiency, with a focus on the societal impacts and safety of AI systems.
The Research Engineer, Pretraining role at Anthropic focuses on developing large language models with an emphasis on safety, alignment, and societal impacts.
The Research Engineer, Performance RL role focuses on advancing AI models' capabilities in safely writing code, collaborating with alignment teams to ensure safety and effectiveness.
The Research Engineer role focuses on advancing the safety and capabilities of large language models through reinforcement learning, collaborating with alignment teams to ensure safe AI systems.
The Research Engineer will redesign how language models interact with external data sources, focusing on safety and societal impacts.
The Research Engineer, Interpretability role at Anthropic focuses on building infrastructure for interpretability research to enhance AI safety through mechanistic understanding of models.
The Research Engineer, Discovery role at Anthropic focuses on developing infrastructure and evaluation frameworks to support the training and deployment of AI systems aimed at achieving scientific AGI.
The Research Engineer will work on advancing AI models in secure coding and vulnerability remediation, blending research and engineering in the field of cybersecurity.
The Model Quality Software Engineer will set technical direction for evaluation systems and research infrastructure, focusing on improving AI model capabilities and ensuring safety in AI systems.
The ML/Research Engineer, Safeguards role at Anthropic focuses on developing systems to detect and mitigate misuse of AI, ensuring safety and compliance in AI systems.