FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.
We’re structured to support that work from early research through real-world adoption:
Independent by design. We can pursue what’s most impactful based on our theory of change and share what we find publicly.
A portfolio approach. Rather than focus on one single direction, we run diverse bets across the safety stack. We take promising ideas from initial experiments to deployment, informed by red-team partnerships with frontier labs and governments.
Serious infrastructure for ambitious research. A dedicated engineering team runs our compute cluster and experiment-scaling stack, so researchers spend their time on research instead of on infra.
Setting the standard. Our events convene key decision makers; our red-team works with frontier developers and governments; and our communications inform the public. Together, this drives adoption and sets the new standard in safety.
Since our founding in July 2022, we’ve grown to 50+ staff, published 40+ academic papers, and convened leading AI safety events. Our work is recognized globally, with publications at premier venues such as NeurIPS, ICML including a Best Paper Honorable Mention in 2026, and ICLR, and features in the Financial Times, Nature News, Wired Magazine and MIT Technology Review. We conduct pre-deployment testing on behalf of frontier developers such as OpenAI and independent evaluations for governments including the EU AI Office and publish the AI Security Leaderboard based on our red-teaming expertise. We help steer and grow the AI safety field through developing research roadmaps with renowned researchers such as Yoshua Bengio; running FAR.Labs, an AI safety-focused co-working space in Berkeley housing 40+ members; and supporting the community through targeted grants to technical researchers.
We explore promising research directions in AI safety and scale up only those showing a high potential for impact. When an approach proves effective, we develop it into a minimum viable demonstration and work with AI developers and governments to support real-world adoption.
Our recent and ongoing research includes:
Adversarial Robustness: working to rigorously solve security problems through building a science of security and robustness for AI, from demonstrating superhuman systems can be vulnerable, to scaling laws for robustness and jailbreaking constitutional classifiers.
Mechanistic Interpretability: finding issues with Sparse Autoencoders, probing deception using AmongUs, understanding learned planning in SokoBan, and interpretable data attribution.
Red-teaming: conducting pre- and post-release adversarial evaluations of frontier models (e.g. Claude 4 Opus, ChatGPT Agent, GPT-5); developing novel attacks to support this work.
Evals: developing evaluations for new threat models, e.g. persuasion and tampering risks, and launching a new research agenda on eval awareness
Mitigating AI deception: studying when lie detectors induce honesty or evasion, and developing approaches to deception and sandbagging.
Applied Interpretability: using interpretability to tackle concrete safety problems (better probes, backdoor detection, deception monitoring), aiming for fast feedback loops, often in collaboration with our other pods.