The ML/Research Engineer, Safeguards role at Anthropic focuses on developing systems to detect and mitigate misuse of AI, ensuring safety and compliance in AI systems.
Curated roles
Browse roles focused on reducing risks from advanced AI systems, including alignment research, safety engineering, evaluations and related operations.
391 active roles found.
The ML/Research Engineer, Safeguards role at Anthropic focuses on developing systems to detect and mitigate misuse of AI, ensuring safety and compliance in AI systems.
The role involves building platforms and tools for reinforcement learning, focusing on data collection, training observability, and ensuring the safety and reliability of AI systems.
The Research Manager for Interpretability at Anthropic will lead a team focused on understanding the internal workings of large language models, emphasizing mechanistic interpretability as a means to enhance AI safety.
The Engineering Manager will lead efforts to improve model performance and ensure the safe development of AI systems at Anthropic.
The Biological Safety Research Scientist at Anthropic will design and develop safety systems for AI, focusing on preventing misuse and ensuring responsible AI safety in the biological domain.
The Applied AI Architect role focuses on guiding enterprise customers in integrating AI systems safely and effectively, ensuring alignment with business objectives and technical implementation.
The Applied AI Architect role at Anthropic focuses on integrating AI solutions into enterprise technology stacks while ensuring safety and reliability, and involves developing evaluation frameworks for AI performance.
The Applied AI Architect role at Anthropic focuses on providing technical guidance to enterprise customers for integrating AI solutions, emphasizing safety and reliability in AI systems.
The Anthropic Fellows Program offers funding and mentorship for research in AI safety and security, focusing on empirical projects aligned with societal benefits.
Research engineer role focused on designing and running model evaluations for Claude, including safety properties, agentic behavior, and evaluation infrastructure at scale.
A full-time AI safety fellowship at Anthropic focused on empirical research projects in AI safety areas such as scalable oversight, adversarial robustness, AI control, and mechanistic interpretability.