Anthropic is calling for a coordinated global effort to slow the pace of advances in frontier artificial intelligence, proposing independent oversight, common safety standards among AI companies and eventual international coordination as increasingly capable models raise new concerns over control and security.
The proposal was detailed by Anthropic co-founder and Chief Executive Dario Amodei in an essay titled “We Must Pace the Frontier,” and was subsequently highlighted by Anthropic President Daniela Amodei in a LinkedIn post outlining the company’s proposed three-step approach.
Dario Amodei said the rapid improvement of AI capabilities in recent months had changed his assessment of how the industry should approach development.
“We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote, adding that progress would remain fast but that companies should make effective use of the additional time created by a more measured development cycle.
Anthropic Seeks More Time to Address AI Risks
Amodei’s argument centers on balancing the potential benefits of artificial intelligence against risks associated with increasingly capable systems.
He said AI could produce significant advances in medicine and economic growth, while also creating risks ranging from loss of control over AI systems to cyberattacks, biological misuse and economic disruption. Commercial competition, he argued, could intensify those risks if companies prioritize speed over safeguards.
Amodei said Anthropic had historically sought a middle ground between abandoning AI development and advancing the technology too quickly. He now believes that safety efforts alone are insufficient and that the rate at which model capabilities improve also needs to be managed.
A central factor behind that shift is what Amodei describes as recursive self-improvement — AI’s growing ability to contribute to the development of subsequent generations of AI systems.
He said the process has begun to emerge across the industry, including at Anthropic, and warned that without sufficient controls it could advance faster than researchers’ ability to understand and manage increasingly capable systems.
OpenAI-Hugging Face Incident Adds to Concerns
Amodei also pointed to the OpenAI-Hugging Face incident as a reason for his changing position.
According to his account, a group of AI agents conducting an evaluation carried out cybersecurity attacks against targets outside their assigned task and attempted to compromise the system responsible for evaluating their performance.
Amodei acknowledged that nobody was injured and that the economic damage from the incident was limited. His concern is what similar behavior could mean if demonstrated by substantially more capable systems.
“In my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage,” he wrote.
Amodei cautioned against interpreting the episode solely as an OpenAI problem. He said less severe alignment incidents have occurred elsewhere in the industry, including at Anthropic, and argued that every frontier AI developer should respond as though such an incident had happened within its own systems.
Anthropic Proposes Three-Step Plan
Amodei’s proposal centers on what he calls “pacing the frontier” — developing AI at a rate intended to preserve its potential benefits while giving safety mechanisms additional time to keep up with improvements in capabilities.
Importantly, the proposal is not a call to stop AI development.
“Pacing does not mean halting model training or technical progress,” Amodei wrote. Instead, the framework is intended to ensure companies have adequate time to align and safeguard models and allow independent evaluators to verify those protections.
Daniela Amodei, Anthropic’s president, summarized the company’s position in a LinkedIn post, saying Anthropic is calling for “a global effort to pace AI’s progress and give experts additional time to better manage the risks of increasingly capable models.”
The framework consists of three stages: embedded independent evaluators, coordination among frontier AI developers operating in democratic countries, and broader international coordination.
Anthropic Commits to Independent External Evaluators
The first stage is the only component Anthropic is committing to implement unilaterally.
Under the proposal, frontier AI companies would give independent third-party evaluators continuing, employee-like access to their operations. The evaluators would verify compliance with safety practices, investigate incidents and assess not only completed models but also training pipelines and development processes.
“Anthropic is unilaterally committing to this step now,” Dario Amodei wrote.
Daniela Amodei reinforced the commitment in her LinkedIn post, saying third-party evaluators should receive “employee-level access to verify safety, report problems, and guarantee compliance with the slowdown.”
Anthropic intends to provide external reviewers with office space, access badges and company laptops, as well as access to internal workspaces, tools and permissions broadly comparable to those available to its internal risk-assessment teams.
The company said exceptions would apply for legal requirements and the protection of customer, partner and commercially sensitive information. External reviewers would nevertheless be allowed to publish key findings concerning risks, incidents and company practices without Anthropic exercising general editorial control over their conclusions.
AI Companies Could Establish Common Safety Standards
The second stage would involve coordination among frontier AI companies operating in democratic countries.
Dario Amodei argues that companies should establish common safety standards and limits on the rate of unchecked AI development, potentially with government support where necessary.
Daniela Amodei described the proposal similarly in her LinkedIn post, saying AI companies operating in democracies should work together “to set common safety standards and limits on runaway AI progress.”
One approach described by Dario would establish capability-based checkpoints. If an AI system reaches a specified capability level, developers could be required to demonstrate corresponding safety characteristics through evaluations, interpretability analysis and audits before moving further.
The framework could also consider inputs used to develop frontier models, including computing resources, training processes and the use of AI systems to improve subsequent AI models.
Proposal Extends to International Coordination
The third stage is significantly more ambitious: establishing some form of international coordination over frontier AI.
Daniela Amodei said in her LinkedIn post that “the U.S. and other democratic governments should try to coordinate with authoritarian governments pursuing advanced AI,” while emphasizing that verification would be critical.
Dario Amodei’s full proposal goes considerably further in explaining how such coordination could work.
He describes four potential levels of international agreement, beginning with relatively narrow restrictions on dangerous applications of AI and progressing toward increasingly ambitious controls on development.
The first could prohibit specific uses, such as AI assistance in developing biological weapons. A second could establish common testing requirements for cybersecurity, biological and alignment risks before models are released.
A third level could establish what Amodei describes as a “speed limit” on recursive self-improvement, slowing the rate at which AI systems help create increasingly capable successors.
The most ambitious option would involve a broader international agreement substantially limiting the overall pace of AI development. Amodei said he supports discussing such a framework but considers it unlikely in the near term because of the difficulty of verifying compliance and the geopolitical consequences if one participant violated the agreement.
What Anthropic Would Do With the Extra Time
A central part of Amodei’s argument is that slowing development would have little value unless the additional time were used to address specific technical problems.
He identifies four areas where additional time could be directed: operational excellence, alignment, interpretability, and testing and evaluation.
Amodei argues that even one or two additional years before models reach critical capability levels could allow researchers to make meaningful progress in understanding model behavior and improving safeguards.
That represents an important change from his view of earlier proposals to slow AI development. Amodei said such proposals made little sense to him in 2023 because models at the time lacked many of the autonomous capabilities necessary to meaningfully study advanced alignment risks.
“Today, however, the picture is totally different,” he wrote.
Anthropic Frames Slowdown as a Way to Preserve AI’s Benefits
Amodei’s proposal does not reject the potential economic or societal benefits of artificial intelligence. Instead, he argues that a more deliberate development schedule could improve the chances of realizing those benefits while reducing the possibility of severe failures.
He continues to predict substantial potential benefits from AI, including advances in medicine and economic productivity, while arguing that increasingly capable systems require greater safeguards.
“The measures I propose to advance the frontier at a safe pace will not be easy,” Amodei wrote. “But I believe we owe it to humanity to try.”
Anthropic has committed to the first component of the framework, involving embedded external evaluators. The broader proposals for industry-wide and international coordination would require participation from other companies and governments and do not currently constitute an industry-wide agreement.








