Anthropic issued a rare public warning this week that its AI system is rapidly approaching the ability to "self-improve"-that is, the ability to independently update its own weights and architecture without human supervision. The company calls on the entire industry to establish a "brake pedal" technical guarantee mechanism that can slow down or stop self-improvement systems that exceed human monitoring capabilities. Anthropic and OpenAI have jointly pressured Congress to establish a mandatory security framework before such systems are deployed.
From "improving while training" to "improving while deploying": A fundamental shift in the risk paradigm
CurrentThe AI security assessment framework is designed on the premise that the model is improved during the training phase (every few months), and then capabilities remain stable during the deployment phase. The self-improvement model breaks this premise:
| dimension | traditional model | Self-improvement model |
| improvement opportunities | Training phase (every 3-6 months) | Deployment phase (ongoing) |
| ability stability | Fixed after release | Continue to evolve after release |
| Effectiveness of safety assessment | Effective evaluation at release | Evaluations quickly become obsolete at release |
| Feasibility of human surveillance | feasible | May be faster than human understanding |
| risk of losing control | Low (predictable) | High (unpredictable) |
Anthropic's specific concern: The current security assessment framework assumes that "models improve between training runs and then maintain stability in capabilities between updates." Self-improving models-systems that can update their own weights or architecture during deployment-represent fundamentally different risk characteristics, because security assessments performed at release may no longer accurately describe the model's capabilities weeks or months later.
"Regulatory capture "charge: rules, but not Sanders 'rules
Anthropic and OpenAI have jointly called on Congress for more safeguards, creating a subtle tension with their opposition to Sanders '50% equity tax bill. Critics point out that this is a classic "regulatory capture" strategy-proactively demanding "rules we can accept" to prevent others from making "rules we cannot accept."
The two companies simultaneously sell to investorsThe trillion-dollar valuation narrative of "models will become stronger" warns the public of the security risk of "models may become too powerful"-this two-sided stance of "both promotion and warning" is particularly sensitive during the IPO season.
Congress's Dilemma: The Time Race between Innovation and Security
The US Congress faces a dilemma:
Acting too fast: May stifle innovation and cede AI leadership to China
Action too slow: Self-improvement systems may be deployed before legislation, creating a "fait accompli"
China's position on this issue is more complicated: on the one hand, ChinaAI companies (DeepSeek, Dark Side of the Moon) are also rapidly advancing model capabilities; on the other hand, China's talent exit restrictions and the "15th Five-Year Plan" show that the government is more concerned about "controllable development" rather than "unrestricted innovation."
Anthropic's "brake pedal" warning is the "Chernobyl moment" in AI safety-not that disaster has already occurred, but that disaster may occur and existing safety systems cannot cope. The deep paradox of this warning is that Anthropic is using the risk that the model may run out of control to service its IPO narrative that the model needs more investment. When a company says both "our technology could destroy humanity" and "please value us at $965 billion," investors need to distinguish between genuine concern or sophisticated marketing strategy.
The deeper problem is that if you improve yourself,AI really appears, is the "brake pedal" enough? The brake pedal of the automobile era can stop a vehicle, but what the "brake pedal" of the AI era needs to stop is a system that may be smarter than humans-it may predict the intention of braking and evade it in advance. Anthropic's warning opened Pandora's Box, but whether there is a real solution in the box remains unknown.