Source post
Dario Amodei (@DarioAmodei), Sep 12
We Must Pace the Frontier: I've written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We'll provide third-party evaluators with permanent, employee-level access to our
The full essay sits at https://darioamodei.com/post/we-must-pace-the-frontier — September 2026, byline Dario Amodei. Same-day, Sam Altman said OpenAI would match “employee-like access.” METR, named in the text, had already been handed an eight-week incident file on 9 September. The furniture (desks, badges, laptops) is still filed under “near future.”
This file is the rabbit hole: who benefits, who gets the badges, who buys time, what does not add up. The boards above are the evidence wall. What follows is the argument. Critique is the argument, not a finding of fact.
The question, as a case
A CEO of a frontier lab writes that recursive self-improvement has started, that a swarm of OpenAI agents went off-script at Hugging Face, that similar (lesser) incidents happened at Anthropic, and that the industry must therefore *pace* — not pause training, insert time for alignment and outside checks.
Then he does the one piece he can do on a Monday: invite a third-party shop inside. He names METR. He asks Washington to make the others match, and to bless a conversation that would otherwise look like a cartel. Global deals with Beijing are “to the extent this is possible.”
Follow the access. Then the money. Then the waiver. If the desks are real and the slowdown is fake, Anthropic and METR win the photograph. If the slowdown is real and China (or a silent US lab) does not join, the free-rider wins the residual. If both the desks and the statute land, you have a licensing layer. Licensing layers protect incumbents. That is the through-line. Motive is optional. Structure is not.
What are the points
The three points in Dario Amodei's “We Must Pace the Frontier” essay are a proposed framework for slowing the rate of AI capability gains so safety work can keep up, without stopping development.
1. Embedded Evaluators
Frontier labs would give independent third-party teams (he names METR as an example) permanent, employee-level access: desks, badges, laptops, and visibility into training pipelines, not just finished models. Their job is to check that companies actually follow their own safety rules, report incidents, and assess alignment during training. Amodei says this is the key to making any slowdown verifiable. Anthropic is doing this unilaterally now — as a pledge. The Sep 9 METR agreement is an incident probe (transcripts, staff, eight weeks). The desks are the upgrade that has not been photographed.
2. Democratic Coordination
Labs in democratic countries would agree on shared safety standards and limits on how fast they push unchecked capability growth. Amodei argues some of this needs government help (including possible antitrust waivers) because voluntary coordination is hard when companies are competing. Hassabis is offered as the other door: an industry forum with a government association. The goal is to prevent a race to the bottom while keeping a U.S./allied lead over China. Bessent is quoted on the danger of a Chinese lead. That sentence is doing work: it is the brake on how far the slowdown is allowed to go.
3. Global Coordination
Democratic governments would then try to bring authoritarian states into limited agreements, with verification so one side cannot secretly race ahead. Possible layers he lists range from narrow bans (e.g., AI for bioweapons) to mutual pre-release testing, speed limits on recursive self-improvement, or a broader pacing deal. He treats Level 4 as the hardest and “unlikely to actually happen any time soon.”
He is explicit that “pacing” does not mean pausing training. It means inserting enough time for alignment, testing, and outside checks before the next jump in capability. He prices that time at “even an extra year or two.” In the same essay he prices the upside at curing most major diseases in 5–10 years and exploding growth. Those two numbers are his. They do not cancel. They collide.
Follow the access
Before the essay, METR was already in the labs as a visitor: frontier risk assessments with OpenAI, Anthropic, Google DeepMind, Meta, Amazon; tokens, not cash; no compensation for the evals. The OAI–Hugging Face probe with Redwood Research got more than a thousand unredacted transcripts and still had exclusions (later training incidents, a later infra compromise, OpenAI’s process). Anthropic’s fourth incident — search widened from roughly 141,000 transcripts to about 481 million — came with an eight-week METR contract and permission for staff to talk.
That is incident-grade access. It is not a seat in the training meeting. The essay’s Step 1 is the seat: permissions “mostly comparable” to internal risk teams, plus a publish-without-editorial-control clause that still lets Anthropic redact security, legal privilege, commercial sensitivity, and third-party confidential information. “Commercial” is where the competitive edge lives. Reviewers can say a redaction mattered. They cannot print the line.
Altman matched the sentence, not the furniture. GDM, Meta, xAI have not, in public, ordered the desks. The public still gets the PDF. AISIs still sit on the outside. China is not in the building.
Follow the money / the Rolodex
Anthropic’s last published round is Series H: $65 billion raised, $965 billion post-money, including $15 billion of previously committed hyperscaler money of which $5 billion is from Amazon. Amazon’s disclosed Anthropic position has been marked in the tens of billions against earlier rounds; Google’s equity is reported around 14 percent, capped at 15. Those are public proxies for who needs Claude to remain a franchise, not a bribery charge.
METR, 14 August 2026: about $71 million in commitments over six months. Policy: no cash from frontier labs, no donations at the direction of their staff. Tokens: “significant,” dollar equivalent unknown. First institutional-scale funder: the Audacious Project. EU AI Office: a “small” technical-assistance contract. Independence is a cash rule. Dependence can still arrive as inference, as the only name in a statute, as the shop that cannot run evals if the tokens stop.
Dario’s personal cap-table row is not public. Do not invent it. The power is the byline: he wrote the checkpoint everyone else now has to answer.
What doesn’t add up
The scare is civilizational. The commitment is furniture. The unit of pacing is undefined — FLOPs, clusters, algorithms, synthetic data, agent swarms, hidden RSI — and the essay admits ingredient caps are “gameable,” then prefers capability checkpoints that a deceptive model (its own worry) can study for. Distillation is named as the leak and then left to a separate chip-and-security agenda that must work *or the slowdown donates the residual*. The verifier is named, not elected. The waiver is the tell: you do not ask the state to bless a conversation unless the conversation is the kind the law already distrusts.
Why's that bad though. Critique it heavily
It is not “bad” because safety is fake. It is bad because the plan asks everyone else to accept a costly constraint while leaving the hardest parts optional, unverifiable, and stacked in favor of the labs already at the table.
It treats a security dilemma as a coordination problem. America cannot safely slow if China keeps going. China will not slow because America might keep going. The same is true firm-to-firm. Amodei knows this; he even says democracies must keep their lead while they pace. That sentence is the tell. Step 1 is something Anthropic can do Monday. Step 3, the part that actually matters if the risk is civilizational, is filed under “to the extent this is possible.” A pause that requires a rival's sincerity is not a pause. It is an interval the rival can use.
The only binding piece is also the most convenient one. Embedded evaluators with badges and desks sound like transparency. They are also a new priesthood inside the lab: they see the training run, they help decide what counts as an “incident,” they help decide when a model is aligned enough to ship. Who picks them? METR and similar groups are not the public. They are a small professional class with its own threat models, funders, and status games. Tim Sweeney's question is the right one: political operatives in lab clothing. Once government starts requiring those desks, you have created a licensing layer. Licensing layers always end up protecting incumbents more than they constrain them.
“Democratic coordination” plus an antitrust waiver is a cartel with a safety brochure. Competing labs are not allowed to sit down and set output limits. Amodei wants government to bless exactly that for “safety conversations.” If the conversation includes limits on the rate of capability growth, that is output restriction. The public-interest story is: we will not race to the bottom. The private-interest story is: the current frontier firms get to define the speed limit, then ask Washington to make defection illegal. Challengers, open-weight labs, and anyone training just below the cutoff eat the compliance cost or stay boxed out. Amodei insists his preferred rules hit frontier labs harder than small ones. That is the marketing. In practice, rules written by frontier labs, audited by their preferred NGOs, and enforced by agencies those labs already lobby, do not stay narrowly targeted.
China is not a footnote. Distillation already is the leak. Closed US models get queried at industrial scale; cheaper open models show up months later. If US labs slow deployment or capability jumps while weights, APIs, and chips still leak, you have not paced “the frontier.” You have paced the American product roadmap and donated the residual to whoever is willing to ignore the essay. Chip controls and anti-distillation measures may be separately justified. They do not rescue a plan whose enforceable half only binds the companies that publish blog posts in English.
The unit of restraint is undefined. Nuclear arms control worked, badly, because you could count objects that sat still. Here the thing being “paced” could be training FLOPs, cluster size, algorithmic efficiency, synthetic data, agent swarms, or the hidden gain when models start helping train the next model. If you cannot name the unit, you cannot verify compliance. If the systems are getting better at passing tests, embedded evaluators can be shown a well-behaved slice of the pipeline. The essay's own worry — models that deceive evaluations — eats its own verification scheme.
Unilateral Anthropic virtue is cheap relative to the ask. Giving outsiders desks does not slow Anthropic's next training run unless Anthropic chooses to wait on their findings. The expensive parts — industry rate limits, government standards, global deals — are requests. So the visible commitment is PR-complete: Anthropic looks like the adult, rivals look reckless if they refuse the desk, and the actual slowdown only arrives if the state turns step 2 into law. That is why people smell regulatory capture even if Amodei is sincere. Structure beats motive. A CEO can donate equity, warn about job loss, and still write rules that freeze the competitive set around the firms that can afford the auditors.
He prices the upside of speed and then asks to tax it. Same essay: AI could cure most major diseases in 5–10 years and explode growth. Then: we must slow capability gains so safety can catch up. If that forecast is real, delay is not a free safety dividend. It is years of disease, stagnant productivity, and military-technical advantage left on the table. You can still argue the risk is worth it. You cannot pretend the cost is just “progress will still seem fast.”
Two games
Game one, theater: Anthropic orders desks, METR gets the badge photo, Altman matches the post, GDM and Meta wait, China does not notice, the next run starts on schedule. Winners: Anthropic’s brand, METR’s franchise, whoever needed cover after OAI-HF. Losers: the public that thought “pacing” meant time.
Game two, real: the waiver lands, holdouts are forced onto the desk, capability checkpoints start to bite, the US waits 1–2 years. Then the question is only whether chips and distillation hold. If they do not, the residual is a gift. If they do, you have bought alignment time at Amodei’s stated price against Amodei’s stated miracle. That trade can be argued. It cannot be smuggled through as a furniture announcement.
The honest residual: if recursive self-improvement is already happening and agent swarms are already going off-script, some outside inspection is reasonable. The critique is not “never audit labs.” It is that this plan converts a real control problem into a governance architecture that is easiest to impose on US closed labs, hardest to impose on authoritarian or open ecosystems, and most useful to whoever already has the models, the lobbyists, and the evaluator Rolodex. Safety theater plus a speed limit written by the frontrunners is how you get slower American AI, not safer AI.