Anthropic issues a "brake pedal" warning-AI's self-improvement capabilities are approaching, and Congress opens an emergency legislative window

6.7 Dimensions Traditional models Self-improvement models Improvement timing Training phase Every 3-6 months Deployment phase Continuous capabilities stability after fixed release Continuous evolution Security assessment effectiveness assessment at release Assessment at release Rapid obsolescence Human surveillance feasibility feasibility may exceed human understanding Speed of out-of-control risk Low predictability High unpredictable

 Anthropic issued a rare public warning this week that its AI system is rapidly approaching the ability to "self-improve"-that is, the ability to independently update its own weights and architecture without human supervision. The company calls on the entire industry to establish a "brake pedal" technical guarantee mechanism that can slow down or stop self-improvement systems that exceed human monitoring capabilities. Anthropic and OpenAI have jointly pressured Congress to establish a mandatory security framework before such systems are deployed.

 


From "improving while training" to "improving while deploying": A fundamental shift in the risk paradigm

 

CurrentThe AI security assessment framework is designed on the premise that the model is improved during the training phase (every few months), and then capabilities remain stable during the deployment phase. The self-improvement model breaks this premise:


dimensiontraditional modelSelf-improvement model
improvement opportunitiesTraining phase (every 3-6 months)Deployment phase (ongoing)
ability stabilityFixed after releaseContinue to evolve after release
Effectiveness of safety assessmentEffective evaluation at releaseEvaluations quickly become obsolete at release
Feasibility of human surveillancefeasibleMay be faster than human understanding
risk of losing controlLow (predictable)High (unpredictable)

 

Anthropic's specific concern: The current security assessment framework assumes that "models improve between training runs and then maintain stability in capabilities between updates." Self-improving models-systems that can update their own weights or architecture during deployment-represent fundamentally different risk characteristics, because security assessments performed at release may no longer accurately describe the model's capabilities weeks or months later.

 


"Regulatory capture "charge: rules, but not Sanders 'rules

 

Anthropic and OpenAI have jointly called on Congress for more safeguards, creating a subtle tension with their opposition to Sanders '50% equity tax bill. Critics point out that this is a classic "regulatory capture" strategy-proactively demanding "rules we can accept" to prevent others from making "rules we cannot accept."

 

The two companies simultaneously sell to investorsThe trillion-dollar valuation narrative of "models will become stronger" warns the public of the security risk of "models may become too powerful"-this two-sided stance of "both promotion and warning" is particularly sensitive during the IPO season.

 


Congress's Dilemma: The Time Race between Innovation and Security

 

The US Congress faces a dilemma:

  • Acting too fast: May stifle innovation and cede AI leadership to China

  • Action too slow: Self-improvement systems may be deployed before legislation, creating a "fait accompli"

 

China's position on this issue is more complicated: on the one hand, ChinaAI companies (DeepSeek, Dark Side of the Moon) are also rapidly advancing model capabilities; on the other hand, China's talent exit restrictions and the "15th Five-Year Plan" show that the government is more concerned about "controllable development" rather than "unrestricted innovation."

 


Anthropic's "brake pedal" warning is the "Chernobyl moment" in AI safety-not that disaster has already occurred, but that disaster may occur and existing safety systems cannot cope. The deep paradox of this warning is that Anthropic is using the risk that the model may run out of control to service its IPO narrative that the model needs more investment. When a company says both "our technology could destroy humanity" and "please value us at $965 billion," investors need to distinguish between genuine concern or sophisticated marketing strategy.

 

The deeper problem is that if you improve yourself,AI really appears, is the "brake pedal" enough? The brake pedal of the automobile era can stop a vehicle, but what the "brake pedal" of the AI era needs to stop is a system that may be smarter than humans-it may predict the intention of braking and evade it in advance. Anthropic's warning opened Pandora's Box, but whether there is a real solution in the box remains unknown.