Research Manager for an AI safety research training programme, supporting teams, technical enablement, proposal evaluation, and participant development.
Curated roles
Browse roles focused on reducing risks from advanced AI systems, including alignment research, safety engineering, evaluations and related operations.
363 active roles found.
Research Manager for an AI safety research training programme, supporting teams, technical enablement, proposal evaluation, and participant development.
Founding Engineer on SaferAI’s Evaluations team, building infrastructure for model evaluations and helping run mitigation-focused safety evaluations for frontier AI labs.
Enforcement analyst role focused on reviewing flagged activity and enforcing policies to prevent misuse of Anthropic's AI systems for cyberattacks, malware, and related harmful operations.
Red team role on Anthropic's Safeguards team focused on adversarial testing of deployed AI systems, model jailbreaking, prompt injection, and detecting novel abuse in frontier AI products.
Part-time technical fellowship building AI evaluation infrastructure, adversarial evaluation tooling, risk data systems, and governance-related prototypes for an AI risk firm.
Project-based adversarial evaluation and red teaming network for frontier AI systems, with work spanning ML security, prompt injection/jailbreaking, agentic system evaluation, and translating findings into governance and mitigation decisions.
Lead threat modeling for CBRNe risks in advanced AI systems, building evaluations and safety frameworks to inform deployment decisions and mitigate dual-use misuse.
Lead role at DeepMind focused on threat modeling and safety evaluations for CBRNe risks in advanced AI models, supporting the Frontier Safety Framework and deployment decisions.
Funding call for foundational research on safety and risk in multi-agent AI systems, including emergent dynamics, trustworthy interaction infrastructure, and scalable monitoring and control.
An 8-week in-person residency for engineers and researchers working on frontier AI security and verification, including projects to mitigate frontier AI risks and secure AI systems and infrastructure.
Research engineer role focused on building and maintaining AI safety evaluation benchmarks, guardrails, and research on agentic failure modes.
Research scientist role focused on AI behavior failure modes, model evaluations, and safety research for LLM agents, including deception, misalignment, unsafe behavior, and frontier agent pressure-testing.
Three AI research engineer roles building the safety-pretraining stack for Apertus, including evaluation/red-teaming, synthetic data, and training systems for open frontier LLMs.
Research engineer role at DeepMind focused on frontier AI safety risk assessment, including measuring and mitigating advanced model risks, loss of control, and harmful manipulation.
Technical program management role leading frontier safety operations, safety frameworks, evaluations, and mitigation processes for Google DeepMind’s frontier AI development.
Research engineer role on DeepMind's Frontier Safety Mitigation team building evaluations, red-teaming, monitoring, and mitigations to reduce misuse and dangerous capabilities in frontier AI models.
Research role on Microsoft’s Futures Team focused on frontier risk and alignment, including frontier safety, normative model behaviour, alignment frameworks, evaluations, and mitigation pathways.
Senior ML engineer role building product experiences, evaluation systems, and trustworthy language-model workflows for research and high-stakes decision-making; relevant because it explicitly emphasizes careful evaluations, process supervision, and safer AI systems.
Role focused on model behavior, prompting, and evaluation pipelines for Perplexity's AI products, including pressure-testing capabilities and validating model behavior before rollouts.
Senior virology and biosecurity role at Google DeepMind focused on biosecurity mitigation, biology evaluations, and protecting frontier AI models from biological misuse.