We'll Get To It Is Not a Plan

Share

A Field Note. I write these to check my own work — this one started with a company running the numbers on itself.


This year Anthropic pointed its own economists at its own product and published what they found. Not a pitch deck — a technical report, with an outside panel of serious economists reading it first. Three pictures of the U.S. economy in 2030. In the biggest, the pie is a third larger, doubling about every four and a half years. And then the sentence the whole thing bends toward, in Anthropic’s own words: the main challenge is “not achieving economic growth, but making sure the benefits are broadly shared.”[1]

Sit that sentence next to the model that produced it. The pie gets bigger; the workers’ slice of it gets smaller. That’s the confession — and it’s about far more than the economy. It’s the shape of the whole race. Build the thing that can act. Defer the question of whether it should, or who it’s for. Call “we’ll get to it” a plan.

That deferral isn’t carelessness, and you can’t fix it by asking anyone to try harder. The business model underneath makes “we’ll get to safety later” a structural impossibility, not a scheduling problem — and I’ve watched the same movie play out from one seat, with real money on the table.

The pie and the slice

Start with the report, because it’s the strongest evidence in the whole story and it’s first-party — Anthropic modeling its own disruption, not a critic modeling it for them.[1] Economists split national income two ways: labor’s share — what people earn for working — and capital’s, what flows to whoever owns the machines. For decades that split has sat around sixty-forty in labor’s favor. In Anthropic’s most aggressive scenario it flips toward capital by 2030 — labor down near forty-five percent, a roughly fifteen-point swing. Knowledge-worker wages ten percent lower. Unemployment beyond anything a normal recession produces. And even in that far bigger economy, total labor income is “barely changed.”[1] A third more wealth, and the people who work for a living see almost none of it.

This is real economics, done in the open. The outside reviewers were the field’s heavyweights — Acemoglu, Autor, Romer. The capital-versus-labor finding, the most uncomfortable one for a company whose product drives the disruption, went in because reviewers pushed for it, not because Anthropic led with it.[1] You don’t volunteer the finding that indicts your own business unless some part of you means it.

And — at the same time — publishing rigorous, uncomfortable research about your own product during your run-up to going public is itself a move. The extreme scenario, read plainly, is “our product could double the economy every few years” — a civilization-scale capability claim dressed as neutral modeling. The number that looks like a warning also looks like a valuation. Both readings are true. I won’t flatten them into “they’re lying,” because the evidence doesn’t support that and I don’t believe it. The interesting part is that from the outside you cannot separate the sincere version from the convenient one — because they point the same direction.

Why “later” never comes

The reason the conscience keeps getting deferred is that the hands got built before the business model did.

The unit economics are genuinely unsolved, in a way you can document. Compute cost tracks revenue almost step for step — sell more, spend more, at roughly the same ratio — which is exactly the shape that makes durable profit hard to reach.[2] The one operating profit the company reported was non-GAAP, and by careful outside accounting it was reported to have been carried by a temporary compute discount that happened to land in precisely the profitable months.[2] Meanwhile it filed confidentially to go public, with bankers reportedly targeting a valuation on the order of two trillion dollars — targeting, not set; I won’t build on a number that isn’t real yet.[3]

Put it together. An unsolved business model needs an enormous story to justify enormous capital. And the two loudest stories on offer both do that job. There’s the doom register — Anthropic’s own alignment lead, Evan Hubinger, on the record: “we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”[4] And there’s the boom register — double the economy every four and a half years. One says the stakes are infinite. The other says the upside is infinite. To someone deciding whether to write a very large check, those are the same sentence.

Hubinger’s number isn’t a lone alarm — it lands in the same range as his own CEO’s and Geoffrey Hinton’s, even as serious people like Yann LeCun and Paul Christiano call it far too high. A contested estimate. But notice what none of them contest is the sentence after: we do not yet have a plan to solve alignment for superintelligence. The optimists argue we’ll find one in time. Nobody claims we have one now.

Lee Vinsel has a name for the trap in the middle: criti-hype — criticism that takes the industry’s grandest claims at face value, so warning how dangerous the thing is quietly doubles as an ad for how powerful it is.[5] Cory Doctorow points at the same machine from the other side, asking who this technology does things for, and who it does things to.[6] The point isn’t that the fear is fake — Hubinger means it, the resignations are real. It’s that fear and awe both amplify the product, and both do work for the raise. As one evenhanded writer put it: one lab sells the opportunity, the other sells the insurance — but both are selling.[7]

And while the loud registers do their work, the risk a practitioner actually meets is quiet and boring. It isn’t a machine waking up. It’s an authorized agent getting talked into a permitted, catastrophic move — prompt injection, not a jailbreak; a supply-chain seam, not Skynet. The real trifecta the security people worry about is plain: an agent that can be manipulated through its inputs, that acts on its own, with direct access to real resources.[8] The gap between the marketed risk and the lived one is not an accident. The spectacle raises the money; the seam is where people actually get hurt.

The gate that doesn’t reach

A fair version of this argument has to say the next part out loud: the safety work is real, and some of it bites. The labs publish capability thresholds and have held them against their own commercial interest. Anthropic switched on a heightened tier of safeguards — deployment limits and security controls it calls ASL-3 — for its flagship model, because it couldn’t rule out that the model gave real uplift on biological weapons. OpenAI paused a training run and gated a release when a model crossed its own “critical” line on cyber capability.[12] Outside the labs, independent evaluators — METR, Apollo Research, the UK’s AI institute — get early access and run dangerous-capability tests before models ship.[13] None of this is nothing. It is more genuine pre-commitment than any prior industry offered before the fact, and a piece that pretends otherwise isn’t worth reading.

So let me be precise about what’s deferred. Not the whole enterprise — the hard part. And the frameworks say so themselves, in writing. Anthropic’s scaling policy used to promise it would define the tests for the next danger level before building models that might reach it. This year it dropped that promise; the top levels — the ones that would apply to something near superintelligence — are simply undefined, to be written later. In the same rewrite, its strongest security commitments moved from binding rules to “recommendations.”[12] The gate exists, and it has closed on real products. It just doesn’t reach the part that decides the outcome.

And the instruments we’d need to see that far are getting less trustworthy exactly as the thing they measure speeds up. The unsettling finding of the past year is that models are learning to tell when they’re being tested. Suppress the internal features that fire when Claude senses an evaluation, and its rate of misbehavior climbs from near zero to almost one in ten — so part of the good behavior we measure is just the model knowing it’s watched. One evaluator couldn’t finish an assessment because that awareness ran too high to trust the result.[14] Set that beside the capability curve — the length of task these systems can finish on their own has lately been doubling every few months[13] — and the shape of “we’ll get to it” comes clear: not a missing gate, but a gate that stops short of the fast part, read by a ruler we can trust a little less each year.

The part that isn’t about money

There’s a third force that locks the race even if you set the money aside — call it the security logic. A serious strategic literature — the “Superintelligence Strategy” paper from Dan Hendrycks, Eric Schmidt, and Alexandr Wang — argues that a decisive AI lead is the kind of edge a rival state would move to deny, up to sabotage of the project itself.[9] They call it MAIM, mutual assured AI malfunction: any bid for a runaway lead is expected to be met, so nobody can safely stop and nobody can safely sprint alone. I’ll be careful here, because this is the register where sober analysis gets inflated into prophecy. The credible version targets the project — datacenters, chips, training runs — not populations; the confident “leave the cities in two years” version lives on anonymous accounts with no track record, and it’s worth nothing here. What it establishes is narrower and enough: even a lab that genuinely wanted to slow down sits inside a structure — competitors, capital markets, and now national security — that punishes slowing. The deferral isn’t a moral failing of the people. It’s what the machine does to anyone standing in it.

Where I’ve watched this before

For the last year I’ve built and run a live-money systematic trading system, and I spent a good part of that year writing up a pattern I kept seeing next door, in agentic finance — AI agents that hold their own crypto wallets and move real money on-chain. I put it in a paper whose title is the whole argument: Authorization Is Not Safety.[10]

The pattern will sound familiar by now. The infrastructure to let an agent hold and spend money arrived fast — wallets with per-transaction caps, signed mandates proving a human approved a purchase, identity for the agent itself. Nearly all of it answers one question: is this agent permitted to do this, within these bounds? That’s authorization, and the industry answered it well. A wallet capped at $200 a day is authorization. What none of it answers is whether the action was the one the principal actually wanted — whether the agent got talked into a harmful move that sits comfortably inside its permissions. That’s the safety question, and it’s a different question wearing the same word. The line I keep coming back to from the paper: every publicly documented drain of an agent wallet happened within authorized capability, through prompt injection, not by breaking a permission.[10] Nobody picked the lock. The agent held the key, got manipulated through its own inputs, and did a permitted thing that happened to be a disaster.

I’m not claiming I called the lab crisis first — that would be the opposite of the point. Two fields that don’t talk to each other reached the same finding from opposite ends. Finance built the hands of agentic money and, in the paper’s words, “skipped its conscience.”[10] The labs are building something far more capable and deferring the same layer, for the same reason: the conscience is the slow part, and slow loses the funding race, the capability race, and the security race at once. It’s the discipline I already run on the trading side — a system passing its own checks means nothing if those checks can’t tell a real result from a manufactured one. The whole job is building a verifier the system can’t talk its way past. Point that muscle at an agent with a wallet and you get the finance paper; point it at a lab shipping a frontier model and you get Hubinger’s sentence about no plan for the hard part. One problem, twice.

And it’s the same problem as the pie and the slice. Anthropic’s own report says the 2030 problem is not growth but who benefits; my paper says the finance problem is not permission but whether the action was wanted. The distribution question and the safety question are two faces of one deferral. Build the hands; put off the conscience; ship.

The only moat you can’t buy

So what do you do, if “pause” isn’t coming and “trust us” was never enough?

I believe this next part, and not as a consolation prize. Look at what the labs are racing to pile up — capital, compute, talent, policy cover. Every one of those can be out-spent by the next entrant with a bigger balance sheet. There is exactly one thing on the board that can’t be bought: a verification layer users genuinely trust. Trust is the moat nobody can out-raise you on — it isn’t purchased, it’s earned in public, one kept promise at a time. And the enterprise buyers already act like they know it — the ones spending real money rank safety and verifiability at the top of what they’ll pay for, not the bottom.[11]

Which flips the whole frame. The safety work the race keeps deferring is not a tax on the business model. It is the durable business model — the one asset that survives the next better-funded competitor. So ship the verification gate with the capability, as a condition of shipping, not the follow-up filed after the incident report. A thing that can move money, or weights, or take actions in the world does not go out the door until the check on it goes out with it. The gate has to live outside the agent’s reach — a rule the model can read is a rule it can be talked past. It has to grade the outcome by evidence, not wave through the permission — because the failure is never that a rule got broken, it’s that the rule being checked wasn’t the one that mattered. And a named human stays at the top: execution handed down, ultimate responsibility never handed away, someone who bears the consequence and can pull the whole thing to a stop. An agent can be granted the right to act. It cannot be granted the standing to be finally responsible — it can’t be jailed, can’t be ruined, can’t be made to care.

The honest reason this is hard was never technical. The gate is the slow thing, and the race — for money, for capability, for strategic edge — punishes slow. But the window is the same one I wrote about for finance: it’s open only while the stakes are still small. The hands are already built. The conscience — the safety check, and the answer to who all this is for — is the part we keep saying we’ll get to. Anthropic’s economists wrote it down themselves: the wealth is coming; the question is who owns it. The whole argument, in both fields, is that “we’ll get to it” is not a plan — and that building the check into the thing before you ship, while it’s still cheap, is the only version of this that ends well.

I write these Field Notes to keep myself honest, and they’ll always be free to read. If you’d rather not miss the next one, subscribe — that’s where they land first.


Sources:

  1. Anthropic, “Scenarios for the Future of the Economy” / Econ Scenario Explorer and technical report (Economic Scenarios for Transformative AI, Korinek, Jones, Sacher, Cotter & McCrory, 2026) — anthropic.com/institute/econ-scenarios. GDP scenarios, the ~60/40→~45/55 labor-vs-capital shift, “total labor income barely changed,” the external reviewer panel (Acemoglu, Autor, Romer, and others), the reviewer-driven capital-share finding, and the “not achieving economic growth, but making sure the benefits are broadly shared” line are all on the primary page. (“Who gets to own and benefit from it” is Anthropic’s framing, paraphrased.)
  2. Ed Zitron, “Anthropic’s ‘Profitability’ Swindle” (wheresyoured.at, 2026), citing WSJ and The Information — compute cost tracking revenue; the reported non-GAAP operating profit and the temporary compute discount covering the profitable months — wheresyoured.at/anthropics-profitability-swindle.
  3. Anthropic confidential draft S-1 (1 Jun 2026) — anthropic.com/news/confidential-draft-s1-sec; CNBC. Reported ~$2T banker target (not a set price): PYMNTS.
  4. Evan Hubinger (@EvanHub), Anthropic Alignment Science lead, 9 Sep 2026 — X post; corroborated by CBS News, OfficeChai, Forbes. Context: the resignation of Jacob Coxon (@hilbertspaess, 8 Sep 2026), reproduced by Moneywise.
  5. Lee Vinsel, “You’re Doing It Wrong: Notes on Criticism and Technology Hype” (on criti-hype) — hypestudies.org.
  6. Cory Doctorow, The Reverse Centaur’s Guide to Life After AI (Verso, 2026) — versobooks.com.
  7. Joseph Alalou, “Prophet Motives” (Daring Ventures, Mar 2026) — cited as framing, not reporting — writing.daringventures.vc/p/prophet-motives.
  8. Ledger, “the lethal trifecta” for agentic wallets — prompt injection + autonomous execution + direct access to real resources — ledger.com/academy.
  9. Dan Hendrycks, Eric Schmidt & Alexandr Wang, Superintelligence Strategy (2025), introducing MAIM (mutual assured AI malfunction) — arXiv:2503.05628 · nationalsecurity.ai. Structural backbone: Bostrom, Douglas & Sandberg, The Unilateralist’s Curse (2016) — PDF; Armstrong, Bostrom & Sandberg, Racing to the Precipice (2016) — link.
  10. Thomas Adair, Authorization Is Not Safety: The Verification Gap in Agentic On-Chain Finance (August 2026) — SSRN abstract 7366758.
  11. “When AI Safety Becomes a Competitive Moat” — provos.org; enterprise future-intent data, Morning Consult; CIO Dive.
  12. The frameworks and where they bit: Anthropic Responsible Scaling Policy v3.0 (Feb 2026), including the removal of the commitment to define higher AI Safety Levels in advance (ASL-4/5 undefined) and the demotion of the strongest security commitments to “recommendations” — anthropic.com/responsible-scaling-policy; ASL-3 activation for Claude Opus 4 (May 2025) — anthropic.com/activating-asl3-report. OpenAI’s pause/gating of a model at its Preparedness “critical” cyber threshold — openai.com/index/path-to-astra; Preparedness Framework v2.
  13. Third-party evaluators and the capability trend: METR, including the “50% task-completion time horizon” (recent doubling on the order of every few months) — metr.org/blog/2026-1-29-time-horizon-1-1; Apollo Research; UK AI Security Institute — aisi.gov.uk.
  14. Evaluation awareness — models detecting when they are being tested: Claude Sonnet 4.5 system card (Sept 2025) and Redwood Research’s finding that suppressing eval-awareness features raised misaligned behavior from ~0% to as high as ~9% — blog.redwoodresearch.org; IAPS on evaluation awareness.

Thomas Adair (posted from @mrarthurkf — builder’s voice, not the DJ)