AI red team engineer role focused on adversary emulation against AI-enabled systems, including the model and its surrounding hardware/software/network stack, to prepare defenders for real-world threats.
Curated roles
Browse roles focused on reducing risks from advanced AI systems, including alignment research, safety engineering, evaluations and related operations.
363 active roles found.
AI red team engineer role focused on adversary emulation against AI-enabled systems, including the model and its surrounding hardware/software/network stack, to prepare defenders for real-world threats.
Cyber-focused AI red team role evaluating model capabilities, safeguards, and abuse risks in agentic systems, with direct responsibility for safety testing and mitigation recommendations.
Senior Model Policy Manager role focused on creating and operationalizing policies, taxonomies, and evaluation criteria to keep frontier AI systems safe, especially for biological and chemical risk.
Founding grantmaker for the AI Model Safety program, funding independent evaluations, standards, and foundational safety research to improve frontier model safety.
Senior AI red team role on SEI's AI Security team focused on adversary emulation against AI-enabled systems, developing offensive techniques to improve defender preparedness and AI system security.
Two-year postdoctoral research role on mechanistic interpretability and AI reasoning within an ERC-funded project on explainable and robust automatic fact checking. Relevant as adjacent AI safety work because it studies model internals and reasoning behavior.
Research scientist role on DeepMind's GenAI Safety team focused on production monitoring, automated evaluations, and misuse detection for deployed AI models.
Senior Model Policy role at OpenAI focused on designing policies, taxonomies, and evaluation criteria to keep frontier models safe in biological and chemical dual-use scenarios.
AI security researcher role focused on LLM security capabilities, benchmarks, risk identification, and security frameworks for model iteration.
Infrastructure/platform engineer building secure, reproducible systems for advanced AI evaluations, including autonomous model behavior, agentic systems, and loss-of-control risk evaluations.
Associate role on Faculty’s AI Safety team supporting frontier model evaluations, AI safety red teaming, and delivery of responsible AI projects for government and industry clients.
Remote full-time GenAI red team role focused on adversarial testing of GenAI safety guardrails, jailbreaks, multi-turn attacks, and model evaluations for cyber-CBRNE threat scenarios.
Remote full-time role focused on evaluating and strengthening GenAI safety guardrails through biological risk assessment, adversarial red-teaming, and model safety alignment.
Infrastructure engineer for Anthropic's Interpretability team, building secure research environments, data systems, and compute tooling that support interpretability work tied to frontier AI safety and audit pipelines.
Senior technical role designing and interpreting evaluations of frontier and specialised AI models for biologically relevant, virology-focused risk and misuse questions, with policy-relevant outputs for government.
Technical research role evaluating risk-relevant capabilities of specialised biological AI models, including misuse risks and technical safeguards, with outputs informing government decisions.
Technical program management role driving AI safety and safeguards initiatives across deployment environments, including model evaluations, mitigations, abuse detection, monitoring, and readiness for high-impact deployments.
A paid MATS research program role involving a ~20 hour AI safety research project focused on pragmatic interpretability or applied safety, with a detailed write-up of findings.
Staff+ SRE role on Anthropic's Safeguards ML Infra team, responsible for production infrastructure, deployment, and validation of safety systems and safety classifiers for Claude model launches.
Research role on OpenAI’s Preparedness team focused on frontier AI safety mitigations, evaluations, red-teaming, and alignment/interpretability methods to make deployed models safer.