Fully-funded PhD position researching responsible machine learning, with explicit focus on AI safety and robust alignment in multi-agent LLM systems.
Curated roles
Browse roles focused on reducing risks from advanced AI systems, including alignment research, safety engineering, evaluations and related operations.
393 active roles found.
Fully-funded PhD position researching responsible machine learning, with explicit focus on AI safety and robust alignment in multi-agent LLM systems.
Research scientist role in Resolution’s philosophy program focused on AI alignment research, including conceptual analysis, empirical hypotheses, and evaluation methods for aligned ASI.
Senior SRE role on Anthropic's Safeguards ML Infra team, focused on deploying, verifying, and operating production safety infrastructure and safety classifiers for frontier model launches.
Program manager for the Canadian AI Safety Institute research program, supporting AI safety initiatives, research partnerships, proposal calls, and related events.
Delivery Manager role in Faculty’s AI Safety team focused on frontier model evaluations, AI safety red teaming, and delivery of high-impact safety projects for clients, including government and frontier labs.
Research role on agent safety, oversight, evaluations, red-teaming, and system-level mitigations for increasingly capable AI agents operating safely and autonomously.
Lead research role at an AI safety organization focused on frontier model risks, misalignment, loss of control, harmful manipulation, and rigorous model evaluation research.
Research manager/generalist for SPAR, an AI safety research fellowship, supporting mentors and mentees, cohort programming, and program operations in the AI safety ecosystem.
Hands-on management role leading Anthropic’s biological safety research engineering team, focused on frontier model evaluations, safety classifiers, red-teaming, and deployment safeguards to prevent catastrophic misuse.
Leads a portfolio of new Kairos programs and ecosystem infrastructure projects aimed at advancing AI safety and reducing risks from advanced AI.
Founding lead for Kairos Labs, an incubator that will design, launch, and support new AI safety organizations and projects.
Lead Kairos’s AI safety university groups portfolio, setting strategy, launching new initiatives, and managing the team supporting AI safety fieldbuilding.
Lead and grow an in-person AI safety residency program for generalists, overseeing recruiting, project matching, advisor coordination, and program strategy for the AI safety ecosystem.
Generalist role supporting Kairos’s AI safety university groups program, including the Pathfinder Fellowship, organizer support, events, and program infrastructure.
Research role focused on mechanistic interpretability and understanding model internals to improve alignment and safety of powerful AI systems.
Analyst role focused on enforcing safeguards against misuse of Anthropic's AI systems for conventional weapons and dangerous technology, including model behavior assessment, enforcement workflows, and evals.
Engineer role focused on cyber-relevant model evaluations, safeguard robustness testing, and misuse detection for frontier AI systems.
Research scientist role on Apollo Research’s AI control and monitoring team, designing control protocols, evaluation frameworks, and monitors to reduce risks from AI systems and coding agents.
Security and control researcher for coding agents, focused on AI threat modeling, red-teaming, and improving monitors and control protocols to reduce catastrophic risks from misaligned or compromised agents.
Dedicated AI red team engineer role focused on red-teaming AI monitors, finding attack surfaces, and improving safety monitoring for coding agents and frontier lab systems.