The White House has asked certain American AI labs, including OpenAI and Anthropic, to reserve access to their latest frontier models for government testing before sharing them with other independent evaluators, such as the AI Security Institute (UK AISI), a leading reference in the field.
- The American administration would thus challenge the privileged pre-deployment access that the British public institute had enjoyed until now.
- The order of access to frontier models would now proceed as follows: first, a government evaluation; then securing American systems; and only afterward sharing the model with foreign partners.
The UK AISI is one of the world’s leading public bodies specializing in pre- and post-deployment evaluation of frontier models.
- Created in November 2023 following the Bletchley Park AI Security Summit, it evaluates most frontier models and documents, among other things, their cyber, biological and chemical capabilities, agent autonomy, and the risks associated with the recursive self-improvement of AI.
- As with other independent evaluators (METR, Redwood Research in particular), the UK AISI could benefit from early access to pre-deployment public checkpoints to run tests on its own benchmarks and document the progression of capabilities and risks associated with AI models.
In early September, the Financial Times reported that the UK AISI had not had access to Mythos 5.1, while it had historically enjoyed early access to major American models. Previously, OpenAI, Anthropic, and Google DeepMind had all provided the Institute with access to some of their most advanced models so they could be tested.
This evolution comes as the cyber capabilities of frontier models become a matter of national security.
- In June, the American administration had already imposed export controls on Fable and Mythos, Anthropic’s models, notably because of the rapid improvement of their cybersecurity capabilities, thus prohibiting their use by foreign nationals.
- The NSA is now testing advanced models to identify vulnerabilities that could affect intelligence and defense systems and is developing its own evaluations. The agency would devote several billions of dollars per year, mainly to computing capabilities.
- The U.S. counterpart to UK AISI, the Center for AI Standards and Innovation (CAISI), remains far more limited in resources than the British institute. Hosted at NIST, CAISI is underfunded and has a smaller research staff.
AI laboratories have also proposed an evaluation directly between labs, without government supervision. They would thereby test each other’s models before market release.
- Google, OpenAI, and Anthropic are jointly working on creating a private standardization body dedicated to AI safety, which could establish common pre-deployment evaluation standards.
- The question of independent evaluation capabilities becomes increasingly central as model capabilities and architectures evolve.
- Redwood Research notes in particular that certain architectures could greatly diminish the ability of human evaluators to monitor and audit model behavior, as they reason not through a comprehensible textual chain of thought but via internal latent states not directly observable.
Beyond tests and evaluation reports, this asymmetry in frontier-model access affects uses in the European industry. The Glasswing project, Anthropic’s early-access program for Mythos, allowed American banks and critical companies to bolster their cyber defense as early as April, notably by using Mythos to probe for vulnerabilities in their own systems. European banks did not gain access until several months later.