devansh · ai made simple · jun–jul 2026

Not a crash. A correction. in one month, every shortcut for trusting AI stopped working

Devansh's frame: June did not prove the artificial intelligence market was fake.

It proved the shortcuts were. The piece.

context intro

The market started asking for receipts

For three years, benchmarks stood in for capability.

Run rates stood in for durable revenue.

Backlogs stood in for future cash flow.

Token prices stood in for cost.

June showed how unreliable those shortcuts had become.

The same question appeared everywhere: how do you verify what you are buying?

This is the correction: the market is discounting claims it cannot check.

01 · the recall

A jailbreak report took two frontier models offline globally

A day after Anthropic shipped Fable 5 and Mythos 5, an outside report showed a prompt trick that made the model surface exploitable vulnerabilities across large codebases.

Commerce issued an export-control directive.

Both models went dark globally because you cannot screen every API user's nationality in real time.

The only enforceable unit was everyone.

Jun 9launch Jun 12dark, worldwide Jun 26~100 vetted US ops Jul 1back, new filter
the part that should scare you

The legal theory was "deemed export": treating a foreign national's API session as a controlled technology transfer.

That theory was never tested in court. The directive just lifted.

Untested means infinitely reusable.

The trigger reportedly ran Amazon → Treasury → Commerce. Amazon is Anthropic's biggest investor, compute vendor, and a competitor.

02 · release becomes permission

The whole industry quietly switched to asking first

Watching that week, everyone changed behavior.

OpenAI launched GPT-5.6 not to the public but to roughly 20 government-coordinated "trusted partners." Google's Gemini 3.5 Pro was suddenly "cleared for July" — a word nobody used two months earlier.

No new law was needed. Memory of the recall did the enforcing.

"voluntary" explicit no-licensing recall risk becomes launch permission 1 · pre-release access = price of launch 2 · state picks the 20 launch partners 3 · the recall — a backstop that never re-fires

Commerce recalled the most capable deployed model on Earth based on a report it could not independently evaluate.

No internal evals.

No severity scale.

No auditor.

That improvisation now forces an audit industry into being. In an audit regime, whoever controls the test controls the market.

03 · the money got audited

The IPO wave dragged the numbers into daylight

Labs filing to go public means one thing: real accounting. Two of this month's tells —

SpaceX IPO, 3 weeks
$135 → ~$225 → below open
open $135 ~$225
TAM claim: $28.5T — a valuation expert called it "written by Grok."
Oracle: booked vs. burned
$638B backlog, −$23.7B free cash flow
$638B backlog −$23.7B free cash
S&P cut Oracle to one notch above junk; its CDS spread went 40 → ~198 bps in a year.

Everyone books backlog: Oracle at $638B, CoreWeave at roughly $99B.

But the GPUs still have to earn it back before financing, depreciation, and rate risk eat the upside.

Demand does not have to vanish. It just has to pay later than the debt assumes.

04 · flat pricing died

Agents broke the one assumption SaaS was built on

Per-seat pricing worked because a human's consumption is bounded. Agents are not.

GitHub moved Copilot to metered "AI Credits" ($0.01 each). Salesforce offered to charge for completed outcomes instead of seats.

Uber blew its whole year's coding-AI budget in four months and capped it at $1,500 per engineer per month. Microsoft dropped most Claude Code licenses even with proof it lifted output — because it owned a close substitute.

same document · same $/token price card… old new +42% 1.42× billable tokens
the tokenizer as a pricing weapon

Claude's new tokenizer produced about 30% more tokens for the same text.

The price card looks unchanged. The effective bill rises once the intro promo ends.

"Dollars per million tokens" is not comparable across vendors when one of them can quietly redefine the token.

05 · open weights took the volume

Cheap open models won the traffic; the frontier kept the margin

Four days after Commerce could switch Anthropic's models off, Z.ai released GLM-5.2 under an MIT license.

Weights you download.

No regional lock.

Nobody can remotely disable them.

That is the open-weight thesis: not ideology, supply-chain control. And the volume followed.

29% of tokens <4% of spend

Chinese models overtook US models in token volume by early June. DeepSeek alone reached 22.6% of token volume.

The market did not replace the frontier. It split the work.

Premium models keep the hard, high-value tasks.

Cheap open models absorb the repetitive layer beneath.

Washington can switch off an American API.

It cannot recall weights already running on private servers.

06 · the frame

Verification is replacing vibes

the one question under all of it

"How do you verify what you're buying?"

Benchmarks can be gamed. Token prices are not comparable.

Routers hide the compute. The top model can still be the wrong product.

The metric that survives: cost per completed task, with the full receipt.

The uncomfortable ending: labs wanted to be the intelligence layer under every industry.

Instead they may become suppliers inside platforms owned by governments, clouds, and enterprise-software companies — because those are the ones who now control the evidence.

zoom out · why it matters now

The month faith stopped being enough

This lands on top of $1T+ in planned AI capex, an IPO wave, and a jittery rate environment.

Whether or not it is a bubble, the market just started demanding receipts.

Every proxy that used to stand in for the truth got audited in the same month.