Remote researcher role producing public reviews of AI benchmarks, evaluating methodologies and implications for AI capabilities; adjacent to AI safety via model evaluations and capability assessment.
Curated roles
Browse remote-friendly roles across AI safety, governance, policy, evaluations, compliance and responsible AI teams.
211 active roles found.
Remote researcher role producing public reviews of AI benchmarks, evaluating methodologies and implications for AI capabilities; adjacent to AI safety via model evaluations and capability assessment.
Remote Futures team role at FLI focused on coalition building, program operations, and communications to advance pro-human AI policy and futures initiatives.
Senior communications leader for an AI safety research organization, responsible for external messaging, media strategy, and translating frontier AI safety research into public-facing materials.
Senior technical role on Cohere’s Safety for Agents team focused on data generation, post-training algorithms, and evaluation methods to improve safety, trustworthiness, and security of LLMs and agentic models.
Research scientist role at Resolution focused on technical AI alignment research, including empirical and theoretical work on scalable oversight and related alignment problems.
Research Engineer role at an ASI alignment lab, building research automation and evaluation infrastructure to support alignment research and frontier-model experiments.
Senior research engineer role on Cohere’s Safety / Modelling Safety and Trust team building data tooling and pipelines for training and evaluation data to support safer, more reliable models.
Research role focused on evaluating frontier AI models on real-world tasks, building benchmarks and rubrics, and analyzing model performance; directly relevant to AI evaluations and safety-adjacent capability assessment.
Technical leadership role building AI verification tools and working with policymakers to support trustworthy adoption and policy change around AI systems.
Research engineer role building AI verification prototypes and supporting technical communications for a verification team whose work is explicitly tied to policy change, international agreements, and trustworthy AI adoption.
Lead research on frontier AI risk modeling across cyber, CBRN, and loss of control, setting methodological standards that inform safety cases, evaluations, mitigations, and AI governance.
Technical Program Manager for Cohere’s public sector and defence AI deployments, with substantial responsibility for governance, compliance, risk assessments, and responsible AI adoption in sensitive government contexts.
Legal counsel role focused on EU AI and privacy regulation for enterprise AI products, including compliance programs, risk assessments, and trustworthy design.
Senior research engineering role focused on building evaluation methods, benchmarks, datasets, and infrastructure for measuring frontier LLM capabilities; relevant because it directly concerns model evaluations for advanced AI systems.
Product management role bridging Cohere's safety research and North product, translating model evaluations and red-teaming findings into safety features, guardrails, and evaluation frameworks.
Senior research role focused on creating next-generation evaluation methods, benchmarks, and infrastructure to measure LLM progress and model capabilities.
Forward Deployed Engineer role building and deploying LLM-powered agentic workflows for enterprise customers, with explicit responsibility for evaluation frameworks, safety, observability, and auditability of AI systems.
Senior technical staff role on Cohere’s Safety for Agents team focused on data generation, post-training algorithms, and evaluation methods to improve safety and trustworthiness of LLMs and agents.
Part-time remote contractor role annotating, auditing, and red-teaming LLM outputs to improve model safety, policy alignment, and prevention of unsafe or adversarial outputs.
Technical program manager for Cohere’s public sector and defence AI deployments, focused on delivery, governance checkpoints, regulatory compliance, and responsible AI considerations.