Product Manager role focused on Anthropic Safeguards systems for frontier models, including safety-by-design, safety evals, detections, interventions, and mitigation of deployment and user risks.
Curated roles
Browse roles focused on reducing risks from advanced AI systems, including alignment research, safety engineering, evaluations and related operations.
393 active roles found.
Product Manager role focused on Anthropic Safeguards systems for frontier models, including safety-by-design, safety evals, detections, interventions, and mitigation of deployment and user risks.
Research fellowship focused on AI alignment and safety research, including alignment challenges in language models and multimodal systems, with work on robustness, interpretability, and evaluation.
Research scientist role on DeepMind’s GenAI Safety team focused on production safety oversight for deployed models, including automated evaluations, misuse detection, and monitoring model alignment and misbehavior.
Product Manager role for Anthropic’s Safeguards team focused on building safety systems, safety evals, detections, and interventions to mitigate misuse and deployment risks for frontier AI products.
Manager role focused on multimodal model safety policy at OpenAI, designing behavioral safety policies, evaluation criteria, and safeguards for frontier AI models.
Program lead for BlueDot Impact’s Technical AI Safety Project Sprint, owning project scoping, participant selection, review, and acceleration for newcomers doing technical AI safety projects.
Research scientist role focused on AI safety, alignment, and model reliability for autonomous driving and ADAS, including safety validation, edge-case evaluation, and robust system behavior.
Penetration tester on Bosch's AI Safety & Security team, performing safety evaluations and red-teaming of LLMs and agentic AI systems to find prompt injection, jailbreak, and tool-abuse issues.
Lead BlueDot Impact’s technical AI safety course, owning curriculum, strategy, admissions, and participant acceleration for a program that trains people entering AI safety.
Senior AI researcher role focused on post-training evaluation, red-teaming, and RL gym audits for open-weight LLMs, with direct work on security capabilities and defensive alignment against indirect prompt injection.
Senior advisory contractor role focused on advanced AI risks, loss-of-control scenarios, safeguards, incident prevention, and translating technical research into policy and operational guidance.
Communications and outreach role at Cedar Research focused on promoting AI safety work through writing, video, outreach, and possibly helping run human studies.
Manager role focused on agentic safety and model policy for frontier AI systems, turning alignment and misalignment risks into behavioral policies, evaluations, monitoring, and safeguards.
Mathematical research role focused on AI safety, including theoretical problems and novel mathematical approaches to safety challenges.
Postdoctoral researcher role on multilingual mechanistic interpretability, including circuit analysis, controlled validation with backdoored model suites, and work that feeds into safe agentic system design.
A remote global funding call for founders to start new organizations tackling critical AI safety problems, including alignment moonshots and frontier capabilities work.
Staff+ software engineer role building production ML infrastructure for Claude's safety systems, including safety deployments, monitoring, and productionizing safety research.
Postdoctoral research role on safe, controllable agentic AI, combining runtime behavioral control, normative constraints, integrated evaluation, and publication in AI safety and formal methods venues.
Technical AI safety research role focused on AI alignment, robust initialization methods for capable language models, and empirical evaluation of alignment techniques.
Research engineer role studying whether values persist after reinforcement learning, using training pipelines, evaluations, and interpretability methods to analyze model behavior and alignment.