We explore promising research directions in AI safety and scale up only those showing a high potential for impact. Once the core research problems are solved, we work to scale them to a minimum viable prototype, demonstrating their validity to AI companies and governments to drive adoption.
Our recent and ongoing research includes:
Adversarial Robustness: working to rigorously solve security problems through building a science of security and robustness for AI, from demonstrating superhuman systems can be vulnerable, to scaling laws for robustness and jailbreaking constitutional classifiers.
Mechanistic Interpretability: finding issues with Sparse Autoencoders, probing deception using AmongUs, understanding learned planning in SokoBan, and interpretable data attribution.
Red-teaming: conducting pre- and post-release adversarial evaluations of frontier models (e.g. Claude 4 Opus, ChatGPT Agent, GPT-5); developing novel attacks to support this work.
Evals: developing evaluations for new threat models, e.g. persuasion and tampering risks.
Mitigating AI deception: studying when lie detectors induce honesty or evasion, and developing approaches to deception and sandbagging.
We are particularly looking to add Research Leads in the following pod shapes:
•
Applied Interpretability — using interpretability to tackle concrete safety problems (better probes, backdoor detection, deception monitoring), aiming for fast feedback loops, often in collaboration with our other pods. A new pod, greenfield.
•
Scalable Oversight / Alignment — methods that keep oversight robust as models become more capable than their supervisors: recursive reward modeling, debate, weak-to-strong generalization, process-based supervision.
•
Adversarial Robustness —extending our independent-testing work into deployed-system protection: better safety guardrails, pre-training safety interventions (initially CBRN misuse, especially for open-weight models), backdoor detection and mitigation, realistic cybersecurity evaluations, and loss-of-control deception evaluations.
•
Auditing / Evals — safety and alignment auditing: evaluation awareness (construct validity, safety-relevance, hyper-realistic evals), CoT monitorability and faithfulness training, black-box monitoring as a complement to our existing white-box work.
•
Persuasion / Epistemic Risks — science of epistemic risks and intervention points, persuasion’s role in loss of control risks, evaluations and independent testing, connections to broader harmful manipulation, solutions and epistemic uplift. Building on our existing work and shaping your own agenda in the area.
•
Bring Your Own Agenda — an open track for senior researchers with a strong vision outside the pods above.
Research Leads define and own a research workstream end-to-end. Day-to-day, that means:
•
Articulate a research agenda with a clear theory of change for mitigating catastrophic risks from human-level or superhuman AI systems, and/or vastly increasing the upside of such systems.
•
Grow and lead a team of technical staff in pursuit of this agenda, either directly or in partnership with an engineering co-lead.
•
Lead novel research projects where there may be unclear markers of progress or success.
•
Share your research findings through written content (e.g. academic publications, blog posts) and presentations (e.g. ML conferences, policymaker briefings) to drive adoption and change.
•
Mentor and coach junior team members in research skills and ML engineering.
•
Contribute to the FAR.AI intellectual environment, for example by giving feedback on early-stage proposals.
•
Build a research field around your agenda through FAR.AI’s grantmaking and events, and connect it to real-world deployments through our independent testing and government advising.
This role would be a great fit if you:
•
Want to work on the most impactful research directions, alongside mission-driven colleagues who’ll push them forward with you.
•
Wish to pursue empirically grounded, scalable research directions that lean, technically strong teams can drive forward.
•
Value the ability to speak freely. We don’t censor our researchers — we just ask that you protect confidential information and make clear when you’re speaking personally or on behalf of the organization.
•
Want to advise and collaborate with governments, leading AI companies, and academics. We’re a small organization that punches above its weight by working closely with these partners — through red-teaming, technical standards work, and research collaborations.
This role would be a poor fit if you:
•
Prefer solo IC research to leading a team toward a shared agenda. Some people can do great research that way, but in this role we’re looking for someone whose research direction is strong enough that other excellent researchers want to build it with them.
•
Prioritize novelty and intellectual elegance over impact. We care about both — a mathematically elegant solution to AI safety would be wonderful — but when we have to choose, we choose what makes AI safer in practice.
•
Can only work with the largest compute clusters available at industry labs or need to be compensated with equity in a rapidly growing startup. We offer competitive salaries and sizable compute budgets on a cluster that we manage, but if you value these things over having a positive impact on the future, then you may be more suited to a for-profit lab.