Research role focused on frontier model evaluations and environments to improve agent capabilities while steering models toward safe AGI/ASI, in collaboration with safety/alignment partners.
Curated roles
Browse roles focused on reducing risks from advanced AI systems, including alignment research, safety engineering, evaluations and related operations.
393 active roles found.
Research role focused on frontier model evaluations and environments to improve agent capabilities while steering models toward safe AGI/ASI, in collaboration with safety/alignment partners.
A full-time, in-person research fellowship in Singapore focused on technical AI safety and governance, with projects aimed at frontier AI safety practice and policy translation.
Editorial role on OpenAI’s Safety Systems team focused on producing and improving public-facing transparency materials about technical safety work, including evaluations, safeguards, red teaming, and deployment decisions for frontier models.
Technical AI Safety Project Sprint BlueDot Impact Remote, Global This course teaches how to make a meaningful contribution to AI safety research or engineering through a structured 30-hour project sprint. Includes weekly mentorship from an AI safety expert to scope, refine, and execute your project with rapid feedback. Publishes projects on the course website as a public signal of your technical AI safety skills to collaborators and employers. Provides collaborative check-ins with peers to receive feedback and build your AI safety portfolio. BlueDot Impact is a nonprofit that runs courses which aim to help participants develop the knowledge, ...
Senior Software Engineer Irregular San Francisco Bay Area In this role, you'll design, build, and scale production systems that power evaluation and security platforms for frontier AI models. Architect and scale production-grade systems and workflows for AI model evaluation. Build backend services, APIs, and monitoring tools that support large-scale evaluations. Design infrastructure that enables research experiments to run efficiently at scale. Implement agent frameworks and build security challenges to test AI models' evasion capabilities. Irregular Labs is a frontier security lab that aims to protect the world in a time of increasingly cap...
The Jailbreaking Lead will focus on identifying and mitigating vulnerabilities in frontier AI models, leading a red team to enhance AI safety and security through hands-on technical work and collaboration with AI developers and governments.
Technical Project Manager, Red Team FAR AI Remote, Global, San Francisco Bay Area, Remote, USA $125,000 - $190,000 In this role, you'll be the delivery backbone of FAR.AI's red-teaming programme, owning engagements with governments and frontier AI companies. Own the end-to-end red-team hiring pipeline, including sourcing, work trials, and recruiting technical talent. Manage the RFP and opportunity pipeline by scoping engagements, drafting proposals, and supporting negotiations. Conduct analysis on team bottlenecks and ecosystem mapping to improve performance. Support technical writing, event organising, policy work, and grant applications as ...
Join Apple's Responsible AI and Safety team as a Senior ML Researcher leading safety alignment and safeguards development for foundation models powering Apple products.
Machine Learning Researcher Gray Swan Remote, USA, Remote, Global In this role, you'll conduct AI safety research to identify and defend against failure modes in advanced AI systems. Discover emerging failure modes through stress-testing cutting-edge AI models. Help enterprises deploy AI safely and at scale without compromising innovation. Inform official safety evaluations of the world's most advanced AI models. Develop research-backed solutions for emerging AI safety challenges. Gray Swan is an AI security company that develops tools that automatically assess the risks of AI models.
Back to careers AI Security Research Engineer Remote / Part-time / Competitive The Mission The last major shift in cybersecurity produced CrowdStrike and SentinelOne. The next shift, AI-driven offense, will be bigger, and defense isn't ready. AI is compressing offensive cyber capability. Attacks that required nation-state resources will soon run autonomously, at machine speed. The industry's response so far has been to replay scripted attack simulations and hope for the best. That's not going to hold. 0Labs is building an AI-native platform and service for continuous purple teaming. Teams of agents execute real, adaptive cyber campaigns, then...
Senior Research Engineer FAR AI San Francisco Bay Area, Remote, Global, Remote, USA $150,000 - $250,000 In this role, you'll accelerate AI safety research by tackling challenging engineering problems and increasing research depth. Lead projects in detecting AI deception, preventing misuse, or building research infrastructure. Mentor team members in technical work to elevate the team's capabilities. Apply software engineering expertise and Python skills to solve complex AI safety challenges. Contribute specialized knowledge in machine learning, high-performance computing, or technical leadership. FAR AI aims to ensure AI systems are trustworth...
Research Scientist Patronus AI Remote, USA, San Francisco Bay Area, New York, NY In this role, you will help solve challenging open research problems facing society’s adoption of AI today, surrounding AI evaluation, language model understanding, and robustness challenges. Develop state-of-the-art systems for AI evaluation and implement algorithms based on NLP advancements. Conduct novel research on redteaming language models, automated evaluation, and alignment. Scope out and lead research projects, including experiment design and understanding results. Patronus AI aims to provide a security and risk management layer for AI by building an aut...
Postdoctoral Researcher, Foundations for Emergent Deception in Human-AI Interaction University of Copenhagen Copenhagen, Denmark In this role, you'll develop computational approaches for understanding emergent deception in human-AI interaction by combining NLP, behavioral psychology, and economic perspectives. Conduct controlled experiments with human participants to study deceptive behavior in AI systems. Apply interpretability-based methods to detect and understand deception in AI. Collaborate with an interdisciplinary team across Computer Science, Psychology, and Economics. Contribute to research advancing frameworks for AI safety and alig...
Join Reflection AI as a Member of Technical Staff - Data Quality Engineer to own upstream data quality and automated QA methods (including LLM-as-a-Judge frameworks) for LLM post-training and evaluation.
Research Engineer Irregular Tel Aviv, Israel In this role, you'll build systems to evaluate and secure frontier AI models at Irregular. Develop infrastructure and experiments to assess model capabilities and implement agent frameworks. Create robust evaluation pipelines and security-focused testing frameworks. Design challenges to measure models' ability to evade detection by defensive security tools. Build controlled environment frameworks and tools that help understand and mitigate risks related to frontier models. Irregular Labs is a frontier security lab that aims to protect the world in a time of increasingly capable and sophisticated AI...
Cyber Researcher Irregular Tel Aviv, Israel In this role, you'll conduct research on AI security, focusing on protecting models against cyber threats and evaluating AI's cybersecurity capabilities. Develop cyber capabilities evaluation challenges including CTF-style vulnerability tests and network attack simulations. Research methods to mitigate AI misuse risks, model weight theft, and dangerous AI agent capabilities. Publish research findings and deliver results to customers. Advise on product development while exploring frontier questions about AI's potential in cybersecurity. Irregular Labs is a frontier security lab that aims to protect t...
Technical Policy Researcher Irregular Tel Aviv, Israel In this role, you'll tackle complex questions of AI development and deployment as a Technical Policy Researcher, shaping the emerging field of AI security. Conduct threat modeling to identify specific ways strong cyber-capable models could cause harm. Develop taxonomies for dangerous AI capabilities and potential mitigations. Create policy proposals for models' refusal policies and track developments in AI policy and research. Write academic papers and blog posts detailing research on evaluation theory and mitigation recommendations. Irregular Labs is a frontier security lab that aims to ...
AI Red Team Analyst Alignerr Remote, Global $15 - $75 per hour In this role, you'll conduct red-teaming exercises to uncover AI security weaknesses and deliver findings that improve system safety. Craft adversarial prompts, jailbreak attempts, and edge-case scenarios to challenge AI model guardrails. Evaluate AI outputs for safety violations, bias, and policy compliance. Document vulnerabilities and unexpected behaviours in structured reports for engineering teams. Collaborate with teams to recommend security mitigations and help refine testing protocols. Alignerr is a platform that for people to create data for frontier AI labs.
Research Manager / Research Managers, AIxCyber ERA Cambridge, UK £62,000 - £75,000 In this role, you'll lead the execution of ERA's AIxCyber Research Fellowship, supporting research fellows in developing and delivering their projects. Review and evaluate fellowship applications, then match fellows with mentors suited to their research focus. Help fellows scope research projects aligned with their skills and interests. Direct and oversee up to 5 research projects while providing guidance on planning, methodology, and publication. Organise workshops and events to support fellow professional development and strengthen relationships across the AI...
Senior communications role focused on external messaging for OpenAI safety research, including alignment, evaluations, preparedness, and interpretability.