Fully-funded PhD position researching responsible machine learning, with explicit focus on AI safety and robust alignment in multi-agent LLM systems.
Curated roles
Browse roles focused on reducing risks from advanced AI systems, including alignment research, safety engineering, evaluations and related operations.
387 active roles found.
Fully-funded PhD position researching responsible machine learning, with explicit focus on AI safety and robust alignment in multi-agent LLM systems.
Senior research role in Oxford’s Technical AI Governance programme focused on interpretability, evaluations, and AI safety for continuously learning systems, with work on foundation-model experiments.
Research engineer role studying whether values persist after reinforcement learning, using training pipelines, evaluations, and interpretability methods to analyze model behavior and alignment.
Mathematical research role focused on AI safety, including theoretical problems and novel mathematical approaches to safety challenges.
Postdoctoral researcher role on multilingual mechanistic interpretability, including circuit analysis, controlled validation with backdoored model suites, and work that feeds into safe agentic system design.
A remote global funding call for founders to start new organizations tackling critical AI safety problems, including alignment moonshots and frontier capabilities work.
Staff+ software engineer role building production ML infrastructure for Claude's safety systems, including safety deployments, monitoring, and productionizing safety research.
Postdoctoral research role on safe, controllable agentic AI, combining runtime behavioral control, normative constraints, integrated evaluation, and publication in AI safety and formal methods venues.
Technical AI safety research role focused on AI alignment, robust initialization methods for capable language models, and empirical evaluation of alignment techniques.
Founding technical research role building white-box auditing methods and infrastructure for frontier AI evaluations, with direct focus on interpretability and safety-relevant failure prediction.
Research scientist role in Resolution’s philosophy program focused on AI alignment research, including conceptual analysis, empirical hypotheses, and evaluation methods for aligned ASI.
Senior SRE role on Anthropic's Safeguards ML Infra team, focused on deploying, verifying, and operating production safety infrastructure and safety classifiers for frontier model launches.
Product design role focused on building and maintaining LLM evaluation systems, prompt fixes, and test harnesses for Claude surfaces and model launches, with explicit safety-related evaluation scope.
Program manager for the Canadian AI Safety Institute research program, supporting AI safety initiatives, research partnerships, proposal calls, and related events.
Senior research role focused on interpretability, evaluations, and AI safety for continuously learning foundation models, with some collaboration on AI governance research.
Delivery Manager role in Faculty’s AI Safety team focused on frontier model evaluations, AI safety red teaming, and delivery of high-impact safety projects for clients, including government and frontier labs.
Research role on agent safety, oversight, evaluations, red-teaming, and system-level mitigations for increasingly capable AI agents operating safely and autonomously.
Lead research role at an AI safety organization focused on frontier model risks, misalignment, loss of control, harmful manipulation, and rigorous model evaluation research.
Research manager/generalist for SPAR, an AI safety research fellowship, supporting mentors and mentees, cohort programming, and program operations in the AI safety ecosystem.
Hands-on management role leading Anthropic’s biological safety research engineering team, focused on frontier model evaluations, safety classifiers, red-teaming, and deployment safeguards to prevent catastrophic misuse.