Join Reflection AI as a Member of Technical Staff - Data Quality Engineer to own upstream data quality and automated QA methods (including LLM-as-a-Judge frameworks) for LLM post-training and evaluation.
Curated roles
Browse roles focused on reducing risks from advanced AI systems, including alignment research, safety engineering, evaluations and related operations.
364 active roles found.
Join Reflection AI as a Member of Technical Staff - Data Quality Engineer to own upstream data quality and automated QA methods (including LLM-as-a-Judge frameworks) for LLM post-training and evaluation.
AI Red Team Analyst Alignerr Remote, Global $15 - $75 per hour In this role, you'll conduct red-teaming exercises to uncover AI security weaknesses and deliver findings that improve system safety. Craft adversarial prompts, jailbreak attempts, and edge-case scenarios to challenge AI model guardrails. Evaluate AI outputs for safety violations, bias, and policy compliance. Document vulnerabilities and unexpected behaviours in structured reports for engineering teams. Collaborate with teams to recommend security mitigations and help refine testing protocols. Alignerr is a platform that for people to create data for frontier AI labs.
Research Engineer Irregular Tel Aviv, Israel In this role, you'll build systems to evaluate and secure frontier AI models at Irregular. Develop infrastructure and experiments to assess model capabilities and implement agent frameworks. Create robust evaluation pipelines and security-focused testing frameworks. Design challenges to measure models' ability to evade detection by defensive security tools. Build controlled environment frameworks and tools that help understand and mitigate risks related to frontier models. Irregular Labs is a frontier security lab that aims to protect the world in a time of increasingly capable and sophisticated AI...
Cyber Researcher Irregular Tel Aviv, Israel In this role, you'll conduct research on AI security, focusing on protecting models against cyber threats and evaluating AI's cybersecurity capabilities. Develop cyber capabilities evaluation challenges including CTF-style vulnerability tests and network attack simulations. Research methods to mitigate AI misuse risks, model weight theft, and dangerous AI agent capabilities. Publish research findings and deliver results to customers. Advise on product development while exploring frontier questions about AI's potential in cybersecurity. Irregular Labs is a frontier security lab that aims to protect t...
Research Scientist Patronus AI Remote, USA, San Francisco Bay Area, New York, NY In this role, you will help solve challenging open research problems facing society’s adoption of AI today, surrounding AI evaluation, language model understanding, and robustness challenges. Develop state-of-the-art systems for AI evaluation and implement algorithms based on NLP advancements. Conduct novel research on redteaming language models, automated evaluation, and alignment. Scope out and lead research projects, including experiment design and understanding results. Patronus AI aims to provide a security and risk management layer for AI by building an aut...
Research Manager / Research Managers, AIxCyber ERA Cambridge, UK £62,000 - £75,000 In this role, you'll lead the execution of ERA's AIxCyber Research Fellowship, supporting research fellows in developing and delivering their projects. Review and evaluate fellowship applications, then match fellows with mentors suited to their research focus. Help fellows scope research projects aligned with their skills and interests. Direct and oversee up to 5 research projects while providing guidance on planning, methodology, and publication. Organise workshops and events to support fellow professional development and strengthen relationships across the AI...
Senior communications role focused on external messaging for OpenAI safety research, including alignment, evaluations, preparedness, and interpretability.
The role involves pursuing a PhD with a focus on AI safety, generative AI, and agentic AI systems, addressing critical aspects of AI development and deployment.
Research engineer building AI platforms and infrastructure for alignment research, safety research, and evaluation of AI alignment/control approaches at CARMA.
The role involves leading the development and implementation of safety systems for AI deployment, focusing on technical safety strategies and evaluations to mitigate risks associated with scientific superintelligence.
The AI Red Teamer role involves probing frontier AI systems for vulnerabilities, designing attack strategies, and building safety evaluation infrastructure to ensure AI safety.
Research Scientist role studying latent structure and behavior in neural networks, with explicit emphasis on understanding as safety, safety-relevant tools, and internal red teaming.
Founding engineer role building an AI-native cyberdefense platform, including AI systems and evaluations that measure performance against real cyber incidents and frontier AI labs.
Technical lead for ARIA’s multi-agent security programme, supporting research on AI agents operating and coordinating in untrusted environments, with explicit emphasis on AI red-teaming and security research.
Computational biologist role focused on building evaluation frameworks and experiments for biological security problems using frontier AI systems.
Role building evaluation task pipelines and infrastructure for frontier model testing, including systems to prevent models from detecting evaluations.
The AI Security & Control Engineer role focuses on designing threat models and control protocols against AI adversaries, improving the security of AI systems, and conducting red-teaming activities to enhance product safety.
Product engineer building interfaces, APIs, and workflows that operationalize interpretability research for training, evaluating, debugging, and deploying AI systems; relevant because it supports safety-related understanding of model internals and safer AI systems.
Funding call for research on AI safety in multi-agent systems, including testbeds, agent networks, infrastructure, and oversight/control for frontier-model agents.
PhD/visiting PhD research role focused on AI safety, security, and alignment for advanced autonomous systems, including interpretability, evaluations, situational awareness, and red-teaming.