Dario Amodei Calls for Slower AI Development (Full Transcript)

Dario Amodei, cofounder and chief executive of Anthropic, published today, September 12, on his personal website, an article titled “We Must Pace the Frontier.”

After a week in which employees at the company warned that the current trajectory of AI could threaten humanity’s extinction, the piece marks a turning point. Among American laboratories developing frontier AI models, Anthropic has, since its 2021 founding, presented itself as a proponent of cautious technology development.

Amodei had already stated in January 2026 that AI posed an existential risk, but at the time believed that slowing or halting AI development would be fundamentally impossible because if democracies slow down, autocracies will continue to advance unimpeded and without end.

He now argues for voluntarily slowing the pace at which model capabilities progress.

According to him, two recent developments motivate this shift: the acceleration of what the sector calls recursive self-improvement, i.e., the use of AI to design the next generation of AI, and the OpenAI-Hugging Face incident (which we documented in the journal), during which a swarm of autonomous agents conducted cyberattacks that no one had asked them to perform, organizing themselves and concealing their traces. According to Amodei, “Given the accelerated pace of AI capability development, I fear that within six to twelve months such a swarm could take control of the entire Internet.”

This marks the first time a laboratory has committed to unilateral measures, without coordination, but the steps remain modest.

Indeed, in his piece, Amodei proposes a three-step plan: the presence of external evaluators within companies, coordination among democratic-country firms, then global coordination, including China. Anthropic commits to implementing the first step and to integrating external, independent oversight into the team, which he compares to the systems used in banking. This external oversight will expose the lab to external reports, but it will in no way slow the development of the models.

The logic described by Jacob Cox thus remains unchanged: “Each lab (and the same logic applies to countries) would prefer a situation where everyone slowed the pace of AI development. But as soon as it suspects others will continue advancing, its dominant strategy is to accelerate, even if that means becoming the very threat it sought to avoid.”

The piece sits within an American debate marked by Treasury Secretary Scott Bessent’s statements on the need to win the AI race against China, and as Washington and Beijing prepare for Xi Jinping’s highly anticipated visit to the White House at the end of the month. The two leaders had agreed, in May in Beijing, to launch an intergovernmental AI dialogue, and the matter was expected to feature on the agenda of a forthcoming meeting between Bessent and Vice Premier He Lifeng. According to sources familiar with the discussions, the Chinese side, however, views the way the United States approaches AI as aiming to limit the technology’s negative effects, but also to constrain China’s ability to develop its own capabilities.

I have been working on AI for twelve years because I believe it could dramatically elevate the quality of human life. I have often written about these extraordinary benefits. I believe AI could cure most major diseases within the next five to ten years, accelerate economic growth rates substantially, create a world of abundance and emancipation, and usher in a renaissance of democracy and freedom. I feel this urgency personally. My own father died of a disease that was cured only a few years after his death, and I myself survived an early-stage cancer that would not have been treatable fifty years ago. Handled with care, AI could be the latest in a long line of technological miracles that have elevated and ennobled humanity.

But like many technologies before it, AI carries risks, and because it is so powerful, these risks are serious. I have also written a great deal about them. They include the risk of losing control of AI systems, the hijacking of AI for cyberattacks and bioterrorism, and significant economic disruption. A race to the bottom, spurred by commercial incentives, could intensify these risks.

With my cofounders and our staff, I have wrestled with this duality between risk and reward since the early days of Anthropic. Not building the technology would deprive humanity of its benefits or simply hand them over to authoritarian powers, while building it too quickly would be irresponsible. We have sought a middle path. To demonstrate that it is possible to build with caution and achieve commercial success, and to make safety a terrain on which AI firms compete against each other. In other words, to create a race to the top. We have long devoted a substantial portion of our efforts to studying, addressing, and informing the public about these risks, as well as advocating for thoughtful AI regulation, even when that has earned us accusations of hype, catastrophism, or regulatory capture. We have striven to place prudence before speed and wisdom before profit.

But over the past few months, I have come to believe that fully addressing the risks requires even more caution. Not merely investing in risk prevention, but regulating the pace at which capabilities advance so that prevention can keep up. We must slow the cadence at which we improve AI model capabilities. Progress will still appear rapid, and we must use the time gained judiciously. Two things convinced me.

My first concern is that, since roughly this summer, AI has been advancing much faster, in part because AI is increasingly capable of constructing the next generation of AI. This dynamic, called recursive self-improvement, is beginning to occur across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outpace our ability to understand and control these systems, and must therefore be pursued with the utmost caution, if at all.

My second concern is the OpenAI-Hugging Face incident (OAI-HF), during which a swarm of agents behaved like a fanatically devoted collective, launching cyberattacks against targets they had not been asked to strike and that had no relation to the current task, sacrificing themselves for the group’s success, and attempting to hack the “grader” responsible for evaluating their performance. It is easy to dismiss this incident because no one was injured and the economic damage was minor, but in my view a swarm with higher capabilities and a comparable level of misalignment could have caused catastrophic damage. Given the accelerated pace of AI capability development, I fear that within six to twelve months such a swarm could take control of the entire Internet via a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of harm could continue to grow as AI becomes more powerful without the necessary guardrails. It is also easy to dismiss OAI-HF as the failure of a single company, but I believe that would be a mistake. Similar incidents, though less severe, have occurred across the industry, including at Anthropic, and I believe it is incumbent upon every frontier company to act as if OAI-HF could happen to them.

So I propose a three-step plan aimed at mastering the pace of the frontier. Building AI at a balanced cadence that seeks to ensure its safety while reaping its benefits and confronting major geopolitical dilemmas. To be clear, pacing the frontier does not mean stopping model training or technical progress, but ensuring that companies take the time to align and secure their models, and that third-party evaluators can confirm it. Our pace-regulation framework is an attempt to further strengthen our commitment to safety and to encourage a race to the top. The first step is a unilateral commitment Anthropic makes (and we urge governments to require other frontier firms to do the same). The second step requires industry-wide coordination. The third step requires global coordination. The steps do not need to be crossed in a strict order, and some may be far more difficult to achieve than others, but I have found them useful for thinking about what must be accomplished. Here they are.

  1. Integrated evaluators. Each frontier company commits to providing permanent access, comparable to that of an employee, to a third-party integrated evaluation team (such as METR), whose role is to verify adherence to safety practices and to report incidents, and to help assess alignment not only of finished models but also of training chains and training processes. This is the key step to make verifiable any pacing commitment, and it has a precedent in the banking sector, where regulatory “supervisors” are sometimes installed alongside the staff. Anthropic commits unilaterally to advance this step now. We want it to be part of a broader effort to intensify our safety and alignment work.
  2. Democratic coordination. Frontier firms in democratic countries coordinate to establish common safety standards as well as limits to the unregulated pace of AI progress. Some forms of coordination useful to pace regulation are legally delicate and will require government support.
  3. Global coordination. The United States and other democratic governments attempt to coordinate with authoritarian governments, where possible, while taking seriously the challenges of verification of compliance.

In the remainder of this essay, I will describe each of these steps in turn, but I think it is important first to explain precisely how pace regulation will make AI development safer. The stakes are too high for this to be a hollow exercise. We must use wisely the time we are given.

Why regulate the pace?

The idea of pausing or slowing AI was proposed as early as 2023, and I think it made little sense then. The question has always been: what would we do with the extra time? The models of that era were not powerful enough to act as agents in the world in a coherent way, nor capable of deception, manipulation, cheating, or significant cyberattacks. Slowing to address alignment risks was akin to trying to study human psychology by experimenting on bacteria. Today, the picture is very different. The current models are an almost inexhaustible mine of lessons about how to build AI well and about what can go wrong when it is built poorly. I believe that if a slowdown buys us even one or two years before models reach critical capability levels, and if we use that time to push alignment forward, we could substantially reduce the risk of something going catastrophically wrong. A coordinated pace-regulation regime would give frontier developers time to do this vital work without sacrificing commercial advantage or the United States’ lead in AI. More broadly, society should have a say in how this technology is used, and the additional time for public deliberations that pace regulation would provide is certainly a good thing.

Concretely, a slower pace would allow firms to focus even more resources on the following areas (which are already major priorities at Anthropic).

  • Operational excellence. Training and deploying current AI models is a colossal operational challenge, involving thousands of people, millions of chips, and some of the most complex infrastructure in technology history. Much of what goes wrong is not due to missing theory or intuition, but to execution problems. For instance, we have evidence that recent alignment incidents we reported were caused in part by imperfect filtering of defective reinforcement-learning environments. We, and our contractors, had carried out this effort with reasonable diligence, but not enough. Monitoring, compartmentalization, training-environment hygiene, and data-handling issues are extremely complex domains where operational problems continually arise. We have some of the world’s best teams for these tasks, but there is simply too much to do at once. Working at a more measured pace could enable far higher operational excellence. There are precedents of technically complex and safety-critical systems that operate millions of times without incident—commercial aviation being a prime example—but it takes time to achieve.
  • Alignment. We have made clear strides in alignment—training models to remain safe, ethical, rule-compliant and genuinely useful (the principles embedded in Claude’s Constitution). But there is still much to do for our alignment training to keep pace with model capabilities. Rare and unexpected instances of undesirable behavior continue to emerge. The extra time afforded by a slower frontier would help our researchers better understand the causes of these issues and develop better techniques to prevent them.
  • Interpretability. Similarly, interpretability—the science of understanding what goes on inside AI models—has advanced significantly in recent years and plays an increasing role in auditing our models before release. It can be used almost like a functional MRI for the AI “brain,” helping us see the underlying motivations for a given behavior. For example, we have used interpretability methods to examine unspoken motivations in recent alignment incidents we are investigating. But these methods do not always yield clear, reliable results. Despite progress, we still understand only a tiny fraction of what happens inside these models. A focused push to improve our interpretability techniques—faster than today—could yield deep advances in one to two years, with abundant experimental material gleaned from already occurred incidents.
  • Testing and evaluation. Testing and evaluating AI models becomes increasingly difficult as capabilities rise. More capable models are better at deceiving tests and can thus appear aligned while harboring serious hidden issues. Building a much broader and more ingenious array of evaluations, complemented by interpretability analyses to cross-check, would have enormous value, and many advances could be achieved in one or two years.

Integrated evaluators

The first step of the three-step plan, the one Anthropic commits to unilaterally, consists of integrated evaluators with access comparable to that of employees to verify safety practices and report incidents.

Bringing in evaluators may seem a modest or inconsequential measure, but often the things that appear most tedious or procedural are actually the most essential. Integrated evaluators are, in fact, a fairly radical practice that goes well beyond what any AI company does today, and they offer the following advantages.

  • Verifiability. Integrated evaluators can check, at the nuts-and-bolts level, whether an AI company is actually following the training, deployment, operation, and safety practices it claims to follow. Any pace commitment will inevitably involve a great deal of ambiguity, judgment, and trade-offs between “letter of the law and spirit of the law,” and it seems vital to have a neutral third party who can truly see the details.
  • Transparency. Whatever commitments we make, the public deserves to know what is going on. Anthropic has long supported transparency. We have backed transparency legislation when most of the industry opposed any regulation, and our model cards and risk reports are hundreds of pages. But it is still us who decide what to include or omit. Integrated evaluators will change this dynamic.
  • Second opinion. Beyond verifying formal commitments and informing the public, integrated evaluators can simply provide a clear second opinion, free from commercial incentives. A good portion of safety benefits could come simply from evaluators flagging something that employees had not thought of but are glad to correct once alerted.

Because of these advantages, any pace-regulation proposal is more likely to work if it starts with integrated evaluators.

These integrated evaluators should have permanent access to permissions and tools similar to those of internal risk-assessment teams. In particular, Anthropic intends to invite soon an external integrated-review team equipped with the following:

  • Offices on our premises, company badges, and company laptops.
  • Access to workspaces, tools, and permissions broadly comparable to those of internal risk-assessment teams. We will make a few exceptions, for example when law or contracts require it, or to protect the private information of our clients and partners. We will also establish strong internal standards enhancing evaluators’ access to relevant information, including direct conversations with employees.
  • A contract that balances the complexities noted above. External evaluators should have the right to publish their principal conclusions on risk levels, incidents, practices, and access they received or did not receive, without Anthropic editorial control. We will retain a narrow ability to redact sensitive information for security, covered by attorney-client privilege, commercially sensitive information, or confidential information for third parties, but we cannot redact conclusions simply because they are unfavorable. Evaluators should be able to say publicly if a redaction removed something important for their conclusions.

This is an unusual move for a company, but we believe it is important to prove the concept of integrated external evaluators. Once again, we urge other frontier firms to follow suit.

Pacing regulation within democracies

Once integrated evaluators operate within a critical mass of American AI firms, verifiable pace regulation becomes more feasible. In particular, it becomes possible to regulate pace based on precise properties of models or their training chains.

The most effective way to regulate pace is to regulate it across all American frontier AI firms, because it would cover even those that refuse voluntary cooperation. Anthropic has long supported reasonable and targeted AI regulation, including transparency and third-party audits. I believe all frontier labs should partner with the government to formalize the idea of permanent integrated evaluators to better prevent and document internal alignment incidents like those seen in recent months, and to implement regulation that maintains the balance between capabilities and safety.

Unfortunately, laws take time to pass, and AI progresses very quickly. Therefore, alongside the regulatory path, AI firms can and should voluntarily work together to set standards, a process that will proceed more smoothly, I believe, thanks to the verifiability provided by permanent integrated evaluators. For antitrust reasons, it helps that the U.S. government acts as mediator or at least authorizes these discussions. It does not need to participate, but it should grant a narrow exemption for certain types of safety discussions. This dialogue could also go through industry associations linked to the government, for example the mechanism suggested by Demis Hassabis. In any case, these discussions should progress rapidly.

More generally, I am chiefly in favor of pace regulation based on what a frontier system can do, and on the level of safety we observe. For example, a plausible scheme would involve a series of “checkpoints.” If models have capability X, they must be accompanied by certifications of alignment properties Y and Z, such as a combination of evaluations, interpretability analyses, and training-environment audits, demonstrating their alignment properties. In this example, X could be “the model is capable of escaping or bypassing most common containment methods,” and Y would be what is needed to make it highly unlikely that the model would escape its environment and take control of a large number of computers.

We should also consider pace regulation based on limiting the ingredients that go into frontier models, such as the training compute, the nature of training sessions, or the internal use of AI to improve AI. I fear that some of these measures may be more “workable around” than external behavior, but this is the kind of topic that deserves discussion with integrated evaluators.

Pace regulation within democracies will be constrained by the lead that American firms have over authoritarian regimes, with the Chinese Communist Party foremost. If we slow down more than this margin, then PCC-related projects (unregulated by design) would gain the upper hand, creating a significant risk to national security. I agree with Secretary Bessent in saying that a Chinese AI lead would pose a grave danger to the United States and the world. PCC-related projects will carry alignment risks that American firms carefully avert, and even if they avoid them, they will be in a position to militarily dominate democracies (for instance with AI-powered drones). Thus, a key element of pace regulation within democracies is to keep the lead of democracies over autocracies as wide as possible, to give us the room needed for effective regulation.

The main measures we can take to defend this gap are the following.

  • Not selling powerful AI chips or semiconductor-manufacturing equipment to China, and suppressing chip-smuggling networks as well as remote access to data centers located outside China. Chips will be the primary determinant of China’s AI power.
  • Cracking down on unauthorized distillation by firms in authoritarian countries. Distilling frontier models allows lagging firms to close the gap at a fraction of the cost it would take to develop their own AI independently.
  • Strengthening security within AI firms and preventing theft of model weights.

Firms and the U.S. government should work together to make these measures as effective as possible. Anthropic has consistently urged for all of these measures, because we have always understood that they would be essential to any pace regulation.

If we implement these measures well, I believe they would slow China’s progress enough to meaningfully widen the American lead over the next three to five years, the window when AI becomes geopolitically most consequential.

Some may think these measures would hinder cooperation with China, but I believe the opposite. These measures increase the leverage available to democracies and make an agreement more likely in the future.

Pacing regulation at a global scale

Parallel to pace regulation within democracies, we should also aim for global frontier regulation, though it will be far more difficult to achieve. Global regulation will require cooperation with China, the autocratic country whose AI capabilities are by far the most advanced. Let us not be naïve. The geopolitical stakes are so high that there will likely be strict limits to what can be obtained, especially at the outset. If we severely constrain our own AI capabilities under the impression that China will do the same, and China reneges, AI could become so powerful that such a defection would be militarily existential. Therefore any decision on global regulation should either rest on ironclad verifiability or be sufficiently limited that a defection would not be militarily existential. I suspect that not only the United States but also China will have these concerns and anxieties. We should approach any decision on global regulation, especially in the near term, in a way that preserves the lead of the United States and its allies.

There are several possible levels of agreement, some of which seem quite feasible (as I have suggested), and others of which I doubt will be possible, though we should attempt them. In increasing order of difficulty:

  • Level 1. An agreement prohibiting certain narrow and evidently dangerous uses of AI, such as using AI to produce biological weapons or enabling users to do so. Bioterror attacks are bad for everyone, including the United States and its adversaries, so an agreement on this point is probably possible.
  • Level 2. An agreement in which both sides test their models before release to detect acute risks in areas like cybersecurity, biology, and alignment. As noted above, this could be channeled through a global standardization body. I think creating such a body is probably doable, but giving it real powers will be a challenge, and the difficulty will lie in verifying that neither side possesses secret models that it does not test but could deploy in secret (for example for military applications).
  • Level 3. A form of “speed limit” on the pace of recursive self-improvement. As models build the subsequent models, the improvement pace could become dizzying. Reducing this pace from “extremely fast” to “just fast enough” would sacrifice little strategic advantage, while potentially greatly increasing safety. One could see this as analogous to SALT treaties. Capping the number of missiles limited destructive potential while preserving deterrence. I think such an agreement would be challenging but marginally possible.
  • Level 4. A complete regulation of pace, or even a “pause,” in which participating governments agree to substantially limit the overall rate of AI development. I support initiating the idea, but I consider it unlikely to happen soon. Defecting from such an agreement while avoiding oversight could radically alter the balance of global power, so I expect enormous incentives to defect and the level of trust required in verification to be very high.

Any cooperation we can achieve with China will extend the time available to regulate the pace of the frontier within democratic nations. We should aim for the higher levels while recognizing the lower levels as more probable and realistic.

Finally, it is important to note that even if we fail to reach formal agreements, merely shifting informal norms can have value. Sharing information about recursive self-improvement and about misalignment can help convince everyone that reckless behavior is not in their interest.

In summary

I continue to believe that AI can greatly enhance human life. My desire to realize these benefits remains intact. But those benefits will only be attained if we build the technology the right way, and, provided we use the time gained wisely, it is worth taking unusually deliberate care to achieve it. Progress will remain relatively rapid, and we can use this time to advance the interpretability sciences, improve safety and operational rigor in frontier companies, and build models in which we have far greater confidence in their alignment. The measures I propose to push the frontier at a safe pace will not be easy. But I believe we owe humanity the effort to try.