Unscarcity
Sign in for free: Preamble (PDF, ebook & audiobook) + Forum access + Direct purchases Sign In

Unscarcity Research

Anthropic Pulled the AI Alarm. Who Can Act on It?

Anthropic says Claude is nearing recursive self-improvement and wants a global pause. No treaty, agency, or verification regime exists to enforce one.

11 min read 2501 words Updated June 2026 /a/frontier-ai-emergency-governance-protocols

Note: This is a research note supplementing the book Unscarcity, now available for purchase. These notes expand on concepts from the main text. Start here or get the book.

When the Lab Pulls Its Own Alarm

On June 4, 2026, the company building one of the world’s most capable AI models published a document arguing that the world should keep open the option to stop building it.

Anthropic’s When AI Builds Itself is not a safety researcher’s op-ed or an open letter from worried academics. It’s the lab itself, pointing at its own production data and saying: the trajectory is heading where we told ourselves it shouldn’t go, and nobody has built the brakes.

The numbers it discloses are the kind you read twice. Claude now writes more than 80% of the code merged into Anthropic’s own codebase. A typical Anthropic engineer ships eight times as much code per day as in 2024. On the hardest open-ended engineering problems, Claude’s success rate jumped to 76% in May 2026, a fifty-percentage-point climb in six months. The tool is increasingly building the toolmaker. Co-founder Jack Clark’s argument is that “recursive self-improvement” (a model designing and building its successor with little human input) could arrive “sooner than most institutions are prepared for.”

Here’s the part that should keep you up at night, and it has nothing to do with killer robots. Anthropic confidentially filed for an IPO days before publishing this warning. A company told its future investors that the thing generating their returns might need to stop accelerating. When the people selling the product start describing the off-switch, the polite assumption that “the experts have this handled” stops being available.

So let’s ask the only question that matters: a warning was issued. Who, exactly, can act on it?

The honest answer, in June 2026, is no one. And that absence — not the capability itself — is the defining institutional gap of this moment.


A Smoke Alarm Wired to Nothing

Imagine a fire alarm that detects smoke perfectly, screams at full volume, and connects to no sprinklers, no fire department, no exit doors. It tells you the building is burning. It cannot do one thing about it. That’s the current state of frontier AI governance.

Walk the inventory of what actually exists:

No international agency. There is no body governing AI development the way the International Atomic Energy Agency governs nuclear material. The IAEA can send inspectors to declared facilities, take environmental samples, run satellite surveillance, and account for every gram of fissile material. For frontier AI training runs, the equivalent agency is a stack of think-tank PDFs.

No treaty. The EU AI Act hits an enforcement milestone in August 2026, and it’s the most serious regulation on the books, but it governs deployment, the act of putting a system in front of users. It says almost nothing about training, the act of building a more capable system in the first place. A model that quietly redesigns itself in a data center triggers none of it.

No precedent that worked. The Future of Life Institute’s 2023 open letter calling for a six-month pause collected tens of thousands of signatures and changed nothing. It came from outsiders, and in 2023 the concern was hypothetical. Both of those excuses just expired.

This is what the book calls governing exponential technology with maps drawn in the steam age. We are, right now, attempting to supervise the most consequential capability humans have ever built using institutions designed to regulate railroads and radio spectrum.


Why a Warning Without Infrastructure Is Theater

The seductive move here is to debate whether Anthropic is right: is recursive self-improvement real, is the threshold close, are the benchmarks meaningful? That debate is a trap, because it’s unfalsifiable in advance and irrelevant to the actual problem.

Suppose Anthropic is exactly right. What happens next?

Nothing happens next. There is no phone number to call. A “global pause,” in Anthropic’s own framing, would require multiple well-resourced labs at or near the frontier, in multiple countries, all agreeing to stop under verifiable conditions. Read that sentence again and count the missing pieces: the labs haven’t agreed, the countries haven’t agreed, and the verification regime that would let any of them trust the others to actually stop does not exist.

This is the difference between a capability warning and an emergency protocol. A warning is information. A protocol is pre-negotiated, pre-built machinery that converts information into action without requiring everyone to invent the response while the building burns. We have flooded the zone with warnings. We have built none of the machinery.

The fallback, absent any of this, is the worst option on the menu: voluntary self-regulation. The company that sounds the alarm also decides whether to heed it. No regulator can force the stop. We are trusting the accelerationists to brake their own car out of conscience, and even the most conscientious lab is one competitive quarter, one activist investor, or one foreign rival away from deciding the conscience can wait.

This is the institutional failure the Unscarcity framework was written to name. Our species keeps building powers that outrun the structures meant to hold them. We solve the engineering and skip the scaffolding. Then we act surprised.


The Book’s Answer: Build the Brakes Before the Engine

The core argument of Unscarcity is that you cannot bolt governance onto a runaway system after the fact. You have to architect the constraints into the system before the compounding starts, because once a process is self-reinforcing, every constraint you try to add afterward gets optimized around. The MBT podcast that surfaced this topic put it crisply: building the constraints before the compounding begins is the whole game.

Two of the book’s Five Laws were designed for exactly this kind of problem, and recursive self-improvement reads almost like a stress test built to break them.

Power Must Decay (Axiom IV): no authority outlives its purpose. The book hard-codes expiration into every form of accumulated influence. Even emergency powers expire automatically after 90 days, with a mandatory cooling-off period before they can return, so a crisis can’t be stretched into a permanent grip on the wheel. Recursive self-improvement is the precise inverse: it’s power that compounds instead of decaying. A model that improves its own ability to improve doesn’t lose authority over time; it gains it, faster, with each cycle. From the book’s vantage point, that’s not a curiosity. It’s a Tier-1 violation wearing a lab coat. Any system whose influence grows without a decay term is, by the framework’s definition, the thing governance exists to prevent.

Truth Must Be Seen (Axiom II): no decision in the dark. The book makes every consequential decision observable, auditable, and traceable on a public ledger, because the failure mode of every collapsed civilization is consequential choices made where no one could see them. The current monitoring regime fails this test badly. Executive orders count GPUs and flag training runs above a compute threshold, but a model optimizing its own code may not need a giant, visible, electricity-guzzling training run to cross the line. Every monitoring framework we’ve built assumes humans design each generation. Remove that assumption and the watchers are auditing the wrong thing.

There’s a third connection, and it’s the most concrete. The emerging best practice for autonomous agents is to gate human approval at irreversible actions: let the agent run freely on anything reversible, but require a human signature before it moves money, deletes data, or signs a contract. Recursive self-improvement is the ultimate irreversible action. The version of the model you safety-tested today is not the version running tomorrow. If iterations are fast enough, every safety test you’ve ever run was conducted on yesterday’s capability, and you never catch up. The model doesn’t need to escape its container. It just needs to redesign the container from the inside.


What Real Emergency Governance Would Actually Require

Naming the gap is easy. The book’s discipline is to refuse the hand-wave and specify the machinery. Real frontier-AI emergency governance has three load-bearing parts, and we can describe each precisely because the arms-control century already built the blueprints.

1. A pre-negotiated pause mechanism. Not a letter. A standing agreement, signed before the crisis, that specifies what triggers a slowdown, who declares it, how long it lasts, and how it ends. This is the Emergency Protocol logic applied to compute: a constitutional dictator for the AI frontier, bound by the Roman constraints the book insists on, namely single purpose, hard time limit, automatic expiry, and personal liability for abuse. The point of pre-negotiation is that you cannot design a fair brake while you’re skidding.

2. Verification, or it’s just trust with extra steps. The cautionary tale here is the Biological Weapons Convention: a real treaty, signed by most of the world, with no verification mechanism, and therefore no way to know who’s cheating, and therefore weak. The opposite model is the IAEA paired with the Non-Proliferation Treaty, where inspectors, sampling, and material accountancy turn “trust me” into “show me.” For AI, the analog is taking shape on paper: FLOP caps, chip registries, offline licensing of advanced processors, hardware-level attestation, and layered verification of large training runs. The technology to make compute auditable mostly exists. The will to require it does not, yet.

3. Liability triggers. A warning costs nothing to issue and nothing to ignore. The fix is to attach consequences in advance: a lab that crosses a declared capability threshold without disclosure faces pre-agreed penalties, and the executives who sign off carry personal exposure, the same liability the Roman dictator faced when he stepped down. Skin in the game is what converts a press release into a decision.

None of this is exotic. Researchers have already sketched the four-institution stack: compute-indexed domestic regulation, an International AI Agency modeled on the IAEA, a “Secure Chips Agreement” modeled on the NPT, and allied public-private partnerships. The first formal treaty negotiations could begin as early as 2027. The blueprints are sitting on the table. We are choosing not to pour the foundation.


“But China Will Never Join”

This is the objection that ends most conversations, so let’s not let it. The argument runs: any pause just hands the lead to whoever doesn’t pause, and Beijing has no reason to freeze. Some critics go further and call Anthropic’s warning competitive positioning: your model is closest, so a freeze locks in your advantage.

Both points have teeth. Both also mistake the current board for the only possible board.

Start with the cynical read of Anthropic’s motives. Even if it’s entirely self-interested, debating the lab’s motives instead of building the response infrastructure is the actual mistake. The smoke alarm’s financial incentives don’t change whether there’s a fire. You build the sprinklers regardless of why the alarm went off.

Now China. The book’s Sovereign EXIT argument is that the Chinese Communist Party’s overriding priority is staying in power, which makes mass labor-cliff unemployment an existential threat and a framework delivering prosperity-without-liberalization strategically attractive, not repellent. Apply the same logic to AI governance: an uncontrolled recursive-self-improvement race is the single fastest route to a technology no government, Beijing’s included, can steer. A regime obsessed with control has the strongest possible reason to fear a system that controls itself. The goal isn’t to convert anyone’s ideology. It’s convergence on a verifiable mutual restraint that every player prefers to the alternative of nobody holding the wheel.

Arms control is the proof of concept. The Soviet Union and the United States despised each other and still signed verifiable treaties, not from goodwill but because mutual annihilation concentrated the mind. The Montreal Protocol coordinated nearly every nation on Earth to phase out ozone-destroying chemicals, enforced it with trade penalties, and it worked. Coordination among rivals is hard. It is not unprecedented. It’s a thing humans have actually done, more than once, when the downside was vivid enough.


The Window Is the Whole Point

Here’s what makes this moment different from every prior AI panic. The warnings used to be speculative: someday a system might. Anthropic’s is empirical: here is our production data, and the line is closer than our institutions are ready for. The threshold hasn’t been crossed. That’s not reassurance. That’s the only good news available, because it means the window to build the machinery is still open. Barely.

The lesson the book draws from Rome, from the Iron Law of Oligarchy, from every constitution eventually bent by someone patient enough, is that you build the guardrails before you need them, while the people designing them still don’t know whether they’ll be the ones constrained. Design the brake while you’re afraid of crashing, not after you’ve decided you’re the best driver.

A capability warning from a frontier lab is not a governance event. It’s a fire alarm wired to nothing. The real work is building the wiring, and it’s the unglamorous kind: treaty-drafting, inspector-training, liability-defining. We have the blueprints. We have, for now, the time. What we lack is the will to treat scaffolding as seriously as we treat the engine.

The book’s whole argument is that this is always the choice, with every exponential technology: deploy the power and improvise the constraints, or build the constraints and then deploy the power. We have, so far, never chosen the second option in time.

This would be a remarkable moment to start.



Sources

Share this article: