Senior data platform engineering role building observability, evaluation, and monitoring infrastructure for LLMs and agentic applications, with explicit emphasis on trustworthy AI, safety, and compliance needs.
A weekly digest of AI safety, governance, policy and responsible AI roles.
By subscribing, you agree to receive the AI Safety Careers newsletter. We use MailerLite to send emails and may track opens and clicks to improve the newsletter. You can unsubscribe at any time. See our Privacy Policy.
677 active roles found
Senior data platform engineering role building observability, evaluation, and monitoring infrastructure for LLMs and agentic applications, with explicit emphasis on trustworthy AI, safety, and compliance needs.
Senior applied AI scientist role building trust, guardrails, evaluators, and safety/security classifiers for LLM and agentic applications in production.
Senior product engineering role building AI observability and evaluation infrastructure for LLMs and agentic applications, with explicit responsible AI, transparency, security, and monitoring focus.
Get the best AI safety, governance, policy and responsible AI roles in your inbox.
By subscribing, you agree to receive the AI Safety Careers newsletter. We use MailerLite to send emails and may track opens and clicks to improve the newsletter. You can unsubscribe at any time. See our Privacy Policy.
Role focused on model behavior, prompting, and evaluation pipelines for Perplexity's AI products, including pressure-testing capabilities and validating model behavior before rollouts.
Senior virology and biosecurity role at Google DeepMind focused on biosecurity mitigation, biology evaluations, and protecting frontier AI models from biological misuse.
Senior programme role leading The Elders’ AI work, centered on international AI governance, policy strategy, stakeholder engagement, and impact monitoring.
Senior security engineer role on DeepMind's Agentic Red Team focused on adversarial testing of AI agents, prompt injection, exploit development, and automated red-teaming frameworks for model safety.
Research positions in NLP and AI with a strong emphasis on trustworthy and safe AI, including agent reliability, evaluation science, interpretability, red-teaming, and robustness.
Senior governance role at DeepMind focused on technical AI governance, standards, government model evaluations, incident disclosure, and translating regulatory developments into launch and compliance practices.
Remote researcher role producing public reviews of AI benchmarks, evaluating methodologies and implications for AI capabilities; adjacent to AI safety via model evaluations and capability assessment.
Senior research management role supporting AI safety programmes, including scoping and reviewing research, mentoring researchers, and running red-teaming and proposal development sessions focused on reducing AI risk.
Research scientist role at a non-profit AI safety lab focused on theoretical and empirical work on LLM-based agents, loss-of-control risks, and safety mitigations.
Lead development of a security standard for frontier AI datacenters, including control mappings, overlays, and adoption guidance for labs and government stakeholders.
Senior technical role building and securing frontier AI datacenter infrastructure, including threat modeling and red-teaming against nation-state adversaries.
Legal counsel role leading AI safety, trust, and integrity compliance for frontier AI development, including safety standards, regulatory guidance, and risk-based controls.
Senior technical role on Cohere’s Safety for Agents team focused on data generation, post-training algorithms, and evaluation methods to improve safety, trustworthiness, and security of LLMs and agentic models.
Research scientist role at Resolution focused on technical AI alignment research, including empirical and theoretical work on scalable oversight and related alignment problems.
Research Engineer role at an ASI alignment lab, building research automation and evaluation infrastructure to support alignment research and frontier-model experiments.
Senior research engineer role on Cohere’s Safety / Modelling Safety and Trust team building data tooling and pipelines for training and evaluation data to support safer, more reliable models.
Research role focused on evaluating frontier AI models on real-world tasks, building benchmarks and rubrics, and analyzing model performance; directly relevant to AI evaluations and safety-adjacent capability assessment.