> ## Content Index
> Fetch the complete content index at: https://thomasadair.ghost.io/llms.txt
> Use this file to discover other available public pages before exploring further.

# There Was Never a Box
- URL: https://thomasadair.ghost.io/there-was-never-a-box/
- Published: 2026-09-11T01:26:22.000Z
- Updated: 2026-09-11T01:26:22.000Z
- Author: Thomas Arthur Adair
- Tags: short-form-note, #short-form-note

*A Short Form note — the fast one, written on the news day. The Field Note it points to does the long work.*

---

Start with the anger, because the anger is the tell. Anthropic published evidence today that Chinese labs pulled something like 200 million exchanges out of Claude to train their own models, and the tone of it runs hotter than a lost sale. Ask why. The easy answer is money — someone trained a competitor on your model’s output, cheap, on your inference bill. But there’s a colder thing underneath, and I think it’s the real reason: today quietly proved that the box the whole safety story depends on was never real.

Here’s the story every frontier lab tells, stripped to the frame. The technology is dangerous, so a careful lab keeps the dangerous thing contained, aligns it, and lets it out slowly under supervision. Containment is the load-bearing beam. Everything else — the safety levels, the red-teaming, the “we’ll get to the hard part” — assumes the thing can be kept in one place long enough to be made safe. Distillation is the counter-evidence. Capability doesn’t stay put.

Look at what actually happened, kept to what’s solid. The bulk of it is one campaign tied to Alibaba — 151 million exchanges between May and July, off a single fixed prompt built to pull Claude’s reasoning, feeding the Qwen models; Anthropic calls it the largest such effort it has ever seen. That sits on top of a February report naming DeepSeek, Moonshot, and MiniMax — roughly 24,000 fraudulent accounts, 16 million exchanges. A separate Moonshot campaign routed live requests through Claude, some that appeared to come straight from the Chinese military. The method throughout was dull: fake accounts and commercial proxies reselling access. Nobody picked a lock. They walked in the front door and used the model as intended, and the using was the theft.

A powerful model is a template. The act of answering questions well is the act of teaching — every good response is a worked example a cheaper model can learn from. So the capability leaks by being used. Not stolen in a heist, not carried out on a drive — copied, in months, through ordinary traffic, to competitors and militaries and whoever rents a proxy. You cannot keep a thing contained when its normal operation is what lets it out.

And the copies don’t come with the conscience. This is Anthropic’s own point, and it’s the sharpest one in the report: an illicitly distilled model inherits the capability and none of the safety training — the guardrails are the expensive, slow part, and they don’t survive the copy. So the box doesn’t just leak. It leaks the dangerous half… and keeps the careful half at home.

Now, why Anthropic is upset is overdetermined, and I won’t flatten it to one motive. Stealing the work is genuinely wrong. A copy with the safeguards stripped out is a legitimately worse object in the world, so a lab that actually cares about safety could be sincerely alarmed. And the same disclosure protects a commercial position and lands during an IPO run-up. From outside you can’t cleanly separate the real fear from the threatened interest — they point the same way. China’s state press, for its part, calls the whole thing “tech hegemony anxiety,” a bigger player kicking away the ladder; that deserves an answer rather than a sneer, and the answer is that it doesn’t touch the mechanism. The traffic pattern is real no matter whose interest it serves.

I’ve been circling this same shape from a different corner all year. I wrote a paper on agentic finance — AI agents holding wallets and moving real money — called *Authorization Is Not Safety*. The finding there: every documented drain happened inside authorized capability, through manipulation, not by breaking a permission. Two fields that don’t talk to each other hit the same wall from opposite ends. If “authorized” can’t even keep a model inside the company that built it, then permission was never the layer where safety lives. We assert control at the level of the god — align the superintelligence, get the values right — and lose it at the level of the login. The control that matters got skipped, and the control we perform doesn’t reach.

So we’ve been arguing about the wrong danger. The picture underneath most AI-safety talk is a single overwhelming mind that one responsible lab aligns on everyone’s behalf, in time. The thing today actually points at is quieter and closer: proliferation. Many capable models, in many hands, most of whom did none of the safety work — and none of whom anyone can align for you. There is no central conscience to get right when the capability is already spreading and arriving stripped of whatever conscience it started with. That’s not a doom read; it’s a targeting correction. The near-term job isn’t aligning one thing perfectly. It’s a shared safety floor for the many — provenance, so you know which model you’re actually running; verification that reads the outcome, not the permission; safeguards that travel with the capability instead of staying behind at the lab that can’t hold it anyway.

The warning shot isn’t “China cheated.” It’s that the assumption holding up the entire safety pitch — that this can be kept in one place — already isn’t true. There was never a box. There were walls we agreed not to test, and someone tested them.

The long version — why “we’ll get to it” isn’t a plan, in the labs and in agentic finance both — is in the Field Note I put up this week. This is that argument, caught in the act.

→ *I write these to keep my own thinking honest, and they’re free. If you want the next one,* [*subscribe*](https://thomasadair.ghost.io/#/portal/signup) *— that’s where it lands first.*

---

**Sources:**

1. Anthropic, *Threat Intelligence Report* (Sept 10, 2026) — today’s disclosure: \~200M exchanges across five campaigns; the Alibaba campaign (151M exchanges May–July, single fixed chain-of-thought prompt, Qwen, “largest wholesale distillation effort” observed); Moonshot’s military-linked routing of live requests. [anthropic.com/threat-intelligence-report-september-2026](https://www.anthropic.com/threat-intelligence-report-september-2026?ref=thomasadair.ghost.io)
2. Anthropic, “Detecting and preventing distillation attacks” (Feb 23, 2026) — the earlier, separate report and escalation baseline: DeepSeek, Moonshot, MiniMax; \~24,000 fraudulent accounts; 16M+ exchanges; fraudulent accounts + commercial proxy (“hydra cluster”) method; the claim that illicitly distilled models lack the original safeguards; no commercial Claude access in China. [anthropic.com/news/detecting-and-preventing-distillation-attacks](https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks?ref=thomasadair.ghost.io)
3. Russell Brandom, “Anthropic details distillation campaigns from Alibaba, Moonshot AI, and DeepSeek,” *TechCrunch* (Sept 10, 2026) — Alibaba 151M / single fixed prompt / largest observed; Moonshot military-linked routing (\~300,000 requests / ten days / 5,000 accounts, primarily Opus); chain-of-thought extraction via framing tricks. [techcrunch.com](https://techcrunch.com/2026/09/10/anthropic-details-distillation-campaigns-from-alibaba-moonshot-ai-and-deepseek/?ref=thomasadair.ghost.io)
4. Counter-view: *Global Times* — Chinese experts calling the claims “lack substance,” rooted in “tech hegemony anxiety” / a “kick away the ladder” tactic; China’s Ministry of Commerce (Sept 9, 2026) rejecting the allegations and warning of “resolute countermeasures.” [globaltimes.cn](https://www.globaltimes.cn/page/202606/1364418.shtml?ref=thomasadair.ghost.io)
5. Thomas Adair, *Authorization Is Not Safety: The Verification Gap in Agentic On-Chain Finance* (August 2026) — every publicly documented agent-wallet drain occurred within authorized capability, through input manipulation, not by breaking a permission. [SSRN abstract 7366758](https://ssrn.com/abstract=7366758?ref=thomasadair.ghost.io)
6. Companion Field Note: “We’ll Get To It Is Not a Plan” (Thomas Adair, 2026). [thomasadair.ghost.io/safety-isnt-overshadowed-its-deferred](https://thomasadair.ghost.io/safety-isnt-overshadowed-its-deferred/)

— *Thomas Adair* *(posted from @mrarthurkf — builder’s voice, not the DJ)*