Research engineer internship focused on AI safety-relevant empirical research on LLMs, including AI security, machine ethics, AI alignment, and benchmarking AI risks.
Curated roles
Browse roles focused on reducing risks from advanced AI systems, including alignment research, safety engineering, evaluations and related operations.
364 active roles found.
Research engineer internship focused on AI safety-relevant empirical research on LLMs, including AI security, machine ethics, AI alignment, and benchmarking AI risks.
Senior operator role owning high-stakes AI safety projects end-to-end, including initiatives like benchmarks for deception and weaponization risk, campaign work on AGI risk, and coordination across researchers and partners.
ML Research Engineer focused on training, post-training, and evaluating LLMs, including alignment methods and safety/moderation-related datasets and policy systems.
AI Red Team Engineer focused on adversarial testing of LLM-powered systems, including jailbreaks, prompt injection, data leakage, policy bypass, and turning findings into regression tests and reports.
Research role building adversarial agent environments, evaluations, and tooling to study AI system failures, misalignment, and unsafe behaviors.
Research engineer role building an AI safety argumentation platform using ontologies, knowledge graphs, defeasible argumentation, and LLM-assisted pipelines to support AI risk management and governance communications.
Contract role focused on monitoring and enforcing abuse on AI products, including building detection, review, and enforcement systems for high-risk harms.
Engineering fellowship supporting AI abuse detection, red teaming, and related research/engineering work across software, data, and ML concentrations.
Mid-level AI red teaming role focused on adversarial testing, jailbreak discovery, failure analysis, and vulnerability reporting for LLMs and image/video models.
Evaluation Engineer role focused on LLM evaluation frameworks, evaluation infrastructure, and production-readiness metrics for enterprise AI systems.
Safety-team role focused on model behavior, alignment, and evaluation of large language models, including building evaluation pipelines and synthetic testing environments.
Offensive security/security researcher role focused on protecting agentic AI systems through AI safety work, red teaming, and model-security risk assessment.
Lead Anthropic’s frontier cyber red team, overseeing research on offensive and defensive capabilities of Claude, model safeguarding, and defenses against advanced AI-enabled cybersecurity risks.
A 3-month full-time AI safety research fellowship with mentorship, reading groups on AI risks, and a final symposium in Cape Town.
Part-time mentor role supervising 3-month research projects for aspiring researchers in AI safety, policy, governance, or biosecurity.
Operations role supporting CAISH’s AI safety programmes, hiring, logistics, and internal systems; relevant because it directly enables AI safety field-building work, though it is not itself a research or policy role.
3-year PhD fellowship researching safety and security evaluation of deployed AI systems, including evaluation, red-teaming, monitoring, and vendor-independent auditing tools for high-stakes deployments.
Security engineer for research infrastructure at an AI alignment nonprofit, building security controls for automated research pipelines, agents, and multi-cloud environments that support frontier alignment research.
Technical communications role translating AI research for policymakers, media, and the public, with explicit responsibility for communicating AI safety and responsible AI research.
Machine Learning Engineer building language-model products, evaluation systems, and trust/transparency features for research and decision-making. Relevant because the role explicitly includes model evaluations and process supervision for more trustworthy AI systems.