We explore promising research directions in AI safety and scale up only those showing a high potential for impact. Once the core research problems are solved, we work to scale them to a minimum viable prototype, demonstrating their validity to AI companies and governments to drive adoption.
We are aiming to rapidly grow our team in the following areas especially, at varying levels of seniority:
•
Evals and red-teaming: Conducting pre- and post-release adversarial evaluations of frontier models (e.g. Claude 4 Opus, ChatGPT Agent, GPT-5); developing novel attacks to support this work; and exploring new threat models (e.g. persuasion, tampering risks).
•
Infrastructure: Maintaining GPU compute infrastructure to support experiments with open-weight models and developing new tooling to allow our research teams to scale their fine-tuning and post-training workflows to frontier open-weight models.
We are also seeking more senior candidates in the following research areas:
•
Mitigating AI deception: Studying when lie detectors induce honesty or evasion, and developing model organisms for deception and sandbagging
•
Adversarial Robustness: Working to rigorously solve these security problems through building a science of security and robustness for AI, from demonstrating superhuman systems can be vulnerable, to scaling laws for robustness and jailbreaking constitutional classifiers
•
Mechanistic Interpretability: Finding issues with Sparse Autoencoders, probing deception using AmongUs, understanding learned planning in SokoBan and interpretable data attribution.
FAR.AI is one of the largest independent AI safety research institutes, and is rapidly growing with the goal of diversifying and deepening our research portfolio. We would welcome the opportunity to add new research directions if you are a senior researcher with a strong vision and would like to pitch us on it.
We organize our team as Members of Technical Staff, with significant overlap between scientist and engineer roles. As a scientist, you will take ownership of and accelerate existing AI alignment research agendas. You can publish research findings broadly and engage with the AI alignment community. If you are an experienced research scientist, then we would be excited to incubate your agenda at FAR using our existing infrastructure and world-class team.
You will receive engineering mentorship via code review, pair programming and regular 1-to-1s. Alongside the engineers, you will be involved in develop scalable implementations of machine learning algorithms and using them to run scientific experiments,
You are encouraged to develop your research taste, proposing novel directions and joining a research pod which suits your interests. You are welcome to take time to study and to attend conferences free of charge. Our technical team is organized into research pods to enable continuity of organizational structure whilst each pod can pivot through varied research projects.
Beyond FAR.AI, you can work with national AI safety institutes, frontier model developers and top academics.