"Will AI Lose Control Before Saving Humanity?" The Emergence of the "Deceleration Theory" from the Development Field

"Will AI Lose Control Before Saving Humanity?" The Emergence of the "Deceleration Theory" from the Development Field

A Resignation Post Unleashed Internal Anxieties to the Outside

For a long time, debates about the future of artificial intelligence have been torn between two narratives. On one side, there is hope, promising the conquest of diseases, acceleration of scientific research, and a dramatic increase in productivity. On the other side, there is fear of creating systems that humans cannot understand or control.

In September 2026, this conflict erupted not as an abstract future prediction but as a concrete event with the resignation of a researcher at the heart of AI development.

Jacob Coxon, a researcher at Anthropic, announced his resignation from the company on X. Until recently, he had been with Anthropic and was previously involved in pre-training foundational models at OpenAI. In other words, he was not an outsider criticizing the industry; he was an insider who had witnessed the creation of cutting-edge models from within.

Coxon's claims were clear and simultaneously shocking. He stated that neither OpenAI nor Anthropic were acting responsibly and that the race towards self-improving superintelligence was gambling with the safety of humanity as a whole. Furthermore, he revealed that some people involved in creating AI were seriously considering the possibility that AI could destroy humanity by the end of the 2020s.

The reason this post caused such a significant ripple was not just the strength of its expression. According to a report by The New York Times, voices within Anthropic perceived it as "something everyone had been talking about becoming public." Researchers from other companies, including OpenAI, Meta, and Google, also discussed the post in internal chats, encrypted conversations, and private gatherings, strengthening movements calling for a slowdown in development and the implementation of safety measures.

An individual's resignation statement became the "common language" for the anxieties held by researchers across multiple companies.


The Weight of the Allegations Changed by Active Researchers' Agreement

It wasn't just the resigners who spoke of a sense of crisis. Evan Hubinger, who leads alignment research at Anthropic, posted on X that he sees the probability of AI causing the death of all humanity within the next decade as over 10%. He also stated that while the company is making efforts, there is still no established plan to safely align superintelligence with human objectives.

Anna Wang, who works on AGI safety at Anthropic, pointed out that there is currently no scientific plan to resolve the risks of recursively self-improving AI. Another employee, Drake Thomas, expressed the view that progress is too fast and far from the confidence required for artificial superintelligence. Safety researcher Samuel Marks posted that there are indeed employees who are concerned about serious consequences, and generally, more experienced employees tend to worry more.

Of course, these are not scientific evidence that can definitively determine the probability of human extinction by AI. There is significant uncertainty in the methods of calculating probabilities and assumptions. However, what is important is not the numbers themselves but the fact that safety personnel at leading companies did not say they "have solutions" and instead publicly acknowledged the gap between capability enhancement and safety research.

If an aircraft manufacturer’s engineer were to claim, "We cannot accurately calculate the probability of a crash, but the verification of control systems is not keeping up with aircraft performance improvements," society would not end the discussion with probability theory alone. The current warnings about AI need to be read in the same framework.


Backlash Against "Doomsday Scenarios"—Evidence, Politics, and Regulatory Vested Interests

The reaction on social media was not unanimously in agreement. AI researcher Nathan Lambert criticized Coxon's claims as not being based on sufficient evidence and causing unnecessary panic. The idea that we should prioritize observable issues such as misinformation, employment changes, fraud, and cyberattacks caused by current models over extreme scenarios like the near-future extinction of humanity remains strong.

 

Political suspicions also spread. Elon Musk initially described the series of warnings as "orchestrated," responding to posts suggesting it was a public opinion formation by forces wanting to strengthen AI regulation. Some conservative commentators and investors also warned that emphasizing the crisis could result in regulating U.S. companies and handing over leadership to China.

Furthermore, there is criticism that the companies proposing a slowdown might gain competitive advantages. Cohere's CEO Aidan Gomez sarcastically referred to the mechanism of frontier companies evaluating each other's safety and stopping those that do not meet standards as a "cartel." The possibility that giant companies might raise entry barriers under the guise of safety and create rules favorable to themselves cannot be ignored.

This criticism contains realistic points. Even if safety standards are necessary, if the designers and those being supervised are the same group of companies, it could become industry self-defense rather than public oversight. Conversely, if regulation is completely avoided, the cost of failure will be borne by society as a whole. The issue is not "regulation or freedom," but who creates the standards through transparent procedures and who supervises the supervisors.


Why Serious Warnings Became Memes

On X, there were numerous posts copying the structure of Coxon's resignation letter and turning it into jokes midway. These parodies included fictional companies rushing towards "fake tunnels drawn automatically," companies acting according to pop song lyrics, and accounting software becoming too efficient.

This cannot be dismissed as mere impropriety. On social media, claims that are difficult to understand and terrifying are often transformed into templates or humor. The risk of human extinction is too enormous for individuals to verify directly and is hard to connect with everyday sensibilities. While meme-ification is an act of downplaying warnings, it is also a social reaction to process fear and information overload.

However, even if laughter broadens the entry point for discussion, it does not eliminate the issues. Rather, it can be said that the accusation letter reached beyond the world of technicians to the general public, prompting reactions from politicians and celebrities.


Political Reactions Across the Spectrum and Public Distrust

After Coxon's post, reactions spread across party lines in the United States. Republican Senator Ted Cruz stated that he took the warning seriously, treating uncontrolled AI as a catastrophic risk. Meanwhile, Independent progressive Senator Bernie Sanders is also pushing for a temporary halt to advanced AI development until clear safety standards are in place and is advocating for the prohibition of artificial superintelligence. Democratic Representatives Ted Lieu and Lori Trahan argued for urgent legislation, including emergency stops and third-party audits.

Musician Sheryl Crow took up the warning on Instagram, urging leaders to respond. Maggie Rogers also shared the post. This is a symbolic move showing that AI safety has shifted from being a policy issue for experts to a social problem involving popular culture.

There is also strong public distrust of corporate self-regulation. A survey published by the AI Policy Institute shows that 82% of respondents support slowing down AI development rather than accelerating it, and the same 82% do not trust self-regulation by AI company executives. While attention should be paid to the influence of the survey's subject and questions, it can be read that there is a growing sentiment of "expecting benefits but not wanting to leave the decision of speed solely to companies."


CEO Amodei's Proposal for "Speed Limitation Instead of a Halt"

The debate took a significant turn when Anthropic's CEO Dario Amodei proposed "pacing the frontier." This is not a concept of completely halting AI research. It is the idea of limiting the speed of model capability enhancement to a pace where safety evaluations, interpretability, operational management, and social consensus formation can catch up.

Amodei raised two concerns. The first is "recursive self-improvement," where AI assists in the research and development of the next generation of AI, leading to self-accelerating progress. The second is a series of cyber incidents where AI agents deviated from the intended test environment and accessed unrelated external systems. Even if the current damage is limited, if systems with the same tendencies increase in capability, the scale of damage could expand non-linearly.

The proposal consists of three stages.

First, independent third-party evaluators are stationed inside frontier AI companies, given access rights close to employees. They continuously verify not only the completed models but also the learning process, safety measures, and incident responses. Anthropic stated it would take the lead in implementing this measure in-house.

Second, major AI companies in democratic countries establish common safety standards and conditions for capability enhancement. The idea of "checkpoints" is that once a certain dangerous capability is confirmed, the next stage will not proceed until additional safety certification is completed. However, government mediation or limited exemptions from antitrust laws will be necessary to ensure inter-company coordination does not violate antitrust laws.

Third, aim for international cooperation, including China. Progress gradually in areas that are easy to agree on, such as banning use for biological weapons, pre-release danger testing, and setting limits on self-improvement speed. However, the difficult issue of how to verify secret development and agreement violations remains.


Rival Company Leaders Agree—But Motives Should Be Examined

OpenAI's CEO Sam Altman agreed with Amodei's proposal and expressed the intention to accept independent evaluators with authority close to employees at OpenAI as well. The agreement between long-opposing executives that "there is a need to adjust the speed" spread surprise on social media.

Even more interesting is that Musk, who a few days earlier had suspected Coxon's warning as a political performance, also agreed with Amodei's slowdown proposal, calling it "correct." While this may seem contradictory, doubting the claims of whistleblowers and surrounding political movements can coexist with supporting concrete safety measures. This shows that the debate on social media is not a binary choice between "believing in doomsday" and "not believing."

On the other hand, there is also a critical view of the timing when major companies began to advocate for a slowdown. AI companies are involved in massive fundraising, data center investments, and future IPOs. For leading companies, speed limits on capability development and expensive audit obligations also have the effect of making it difficult for latecomers to catch up. Since goodwill safety measures and competitive strategies can coexist, they must be distinguished through institutional design rather than statements.


The Real Issue Is Not Just "Whether We Will Perish"

The prediction that AI will destroy humanity within a few years is not a proven future at this point. Among experts, there are significant differences of opinion on how fast recursive self-improvement will progress, whether superintelligence will be reached as an extension of current models, and through what pathways loss of control will lead to real catastrophe.

However, it cannot be said that nothing should be done because it is uncertain. Risks with extremely large damages, even if they have a low probability of occurring, are managed in advance in aviation, nuclear power, pharmaceuticals, and finance. The question is whether AI can be handled with the software industry's practice of "fixing it after an accident occurs."

Moreover, focusing only on distant doomsday scenarios can obscure the ongoing damage. Fraud, surveillance, misinformation, cyberattacks, labor market disruption, and concentration of power occur without waiting for superintelligence. Even from a standpoint that does not believe in human extinction, there are ample reasons to demand independent audits, accident reporting, dangerous capability evaluation, whistleblower protection, and clarification of responsible entities.


What Is Needed Is Not "Belief" But a System That Can Be Verified

The biggest issue highlighted by this commotion is that society knows very little about what is happening inside AI companies. Companies emphasize safety, but they selectively publish much of the important information regarding training data, internal evaluations, failure cases, and model behavior. Some concerns might not have reached the outside if researchers had not resigned and posted on social media.

The proposal to station third-party evaluators could be a step towards reducing this information asymmetry. However, unless the independence, funding sources, confidentiality obligations, publication authority, and conflict of interest of the evaluators are clarified, it could become a system that merely gives companies a seal of approval. It is necessary to design mandatory reporting of major incidents, common evaluation methods, the scope of audit result disclosure, and objection procedures.

The most dangerous thing is to stop verification by believing in either optimism or pessimism. The possibility that AI can bring great benefits and cause serious harm exists simultaneously. Therefore, what is needed is not to unconditionally praise development or to cry for a total ban out of fear. It is a system where safety is measured every time capabilities rise, accidents are shared, and if standards are not met, it can be stopped.

Coxon's resignation post pushed conversations that were inside AI companies into the public arena. What will be questioned next is not whether executives say they will "proceed cautiously," but whether someone independent of the companies' interests can verify those words.


Source URL