Skip to main content

The AI price war is here: how a 24-hour launch blitz cut token costs ~80%

6 min read

In one 24-hour window, three major AI models launched and output-token prices collapsed from $25–50 to $4–6 per million. Here's what the price war means for your bottom line.

The AI price war is here: how a 24-hour launch blitz cut token costs ~80%

Picture this: you go to bed on July 8, and by the time you wake up the next morning, three of the biggest names in AI have shipped new flagship models — and the cost of running them has fallen off a cliff. That’s not a hypothetical. It happened, and it quietly rewrote the math for anyone building with AI.

What actually happened: a 24-hour blitz

Here’s the sequence. On July 8, 2026, xAI dropped Grok 4.5. Less than a day later, OpenAI pushed its GPT-5.6 family to general availability, and Meta shipped Muse Spark 1.1. Three labs, one 24-hour window, all racing on the same thing: making a capable model that’s fast and cheap.

That timing wasn’t a coincidence — it was a shot fired in a price war that’s been building all year. When rivals launch within hours of each other, nobody wants to be the expensive option on the shelf.

The 24-hour blitz

Three labs, one day. Each shipped a flagship built to be fast and cheap. Hover a card.

Jul 8 · xAI
Grok 4.5
Coding & agents
$2 in / $6 out
Uses ~¼ the output tokens of legacy flagships per solved task.
Jul 9 · OpenAI
GPT-5.6 Luna
High-volume tier
$1 in / $6 out
The cheap, fast member of the new Sol / Terra / Luna lineup.
Jul 9 · Meta
Muse Spark 1.1
Agentic work
1M context
Leaned hard into agents and computer use with a 1M-token window.

The number that actually matters

Ignore the benchmark bragging for a second. For anyone paying the bill, the story is output tokens — the text a model generates, which costs several times more than the text you feed in. That’s where your invoice really comes from.

And output pricing just cratered. Legacy flagship models sat in the $25–50 per million range. The new wave? Grok 4.5 and GPT-5.6 Luna both land at $6, with Gemini 3.6 Flash at $7.50. That’s roughly an 80% cut for output that’s often just as capable for everyday work.

The output-token price crash (per 1M tokens)

Output tokens are where the bill lands. Watch the newest models undercut the legacy flagships. Hover a bar.

Claude Fable 5 (legacy)
$50
GPT-5.6 Sol
$30
GPT-5.6 Terra
$15
Kimi K3
$15
Gemini 3.6 Flash
$7.50
Grok 4.5
$6
GPT-5.6 Luna
$6
Legacy / flagship tierMid tierNew cheap wave

What that means in real money

Percentages are abstract, so let’s make it concrete. Take a fixed budget — say $100 of output — and see how far it stretches at each tier.

What $100 of output tokens buys you now

Same $100, very different output. The new wave stretches your budget ~8× further than a legacy flagship.

Legacy flagship ($50/M)
2.0M tokens
Mid tier ($15/M)
6.7M tokens
New wave ($6/M)
16.7M tokens

At old flagship rates ($50/M), $100 buys you 2 million output tokens. On the new wave ($6/M), the same $100 buys nearly 17 million — about 8× more work for the same spend. For a content operation, an agent that runs all day, or a coding workflow, that’s not a rounding error. That’s the difference between a side project that bleeds money and one that turns a profit.

OpenAI’s three-way split

The clearest sign of the strategy is what OpenAI did with GPT-5.6: it split it into three tiers instead of one model. Sol ($5/$30) is the flagship for the hardest reasoning and agentic work. Terra ($2.50/$15) matches the old GPT-5.5 at roughly half the cost. And Luna ($1/$6) is a brand-new cheap tier built for high-volume jobs.

Notice the pattern — it’s the same “workhorse, balanced, budget” split Google used with Gemini 3.6 Flash and Flash-Lite. The labs have realized that most of us don’t need the frontier model for most tasks; we need a cheap-but-good one, and a premium one for the hard 10%.

And the floor keeps dropping

Think $6 is cheap? The bottom is lower still. Among the Chinese open models, DeepSeek V3.1 Terminus sits at just $0.27 per million input — the cheapest flagship-class model tracked. Kimi K3, the 2.8-trillion-parameter open giant, runs $3/$15. The open-weight camp keeps yanking the price floor down, and the closed labs have to answer.

Here’s the thing: none of this is slowing down. Cheap, capable AI is becoming a commodity, and commodities compete on price. Great news if you’re the buyer.

My builder’s read: how to actually profit from this

A price war only helps you if you act on it. Three concrete moves:

  • Re-price your stack this quarter. Take your two or three heaviest recurring tasks and check what they’d cost on the new-wave models versus what you’re paying now. You’ll almost certainly find a task you can move for 70–80% less at equal quality.
  • Route by job, not by loyalty. Send high-volume, low-stakes work to a $6 model; reserve the pricey flagship for the genuinely hard 10%. That’s the whole multi-model habit — and it’s never paid off more than right now.
  • Test before you switch. Run five real tasks from your own workload through a cheap contender and your current default, and score quality, speed, and cost. Let your work decide, not the launch thread.

The catch nobody’s putting in the headline

Now the honest part. Cheaper tokens are a discount only if your usage stays flat. The trap is subtle: when output drops to $6, it’s tempting to 10× how much you generate — more drafts, longer agent runs, chattier prompts — and suddenly your “cheaper” bill is bigger than before.

So watch your total monthly spend, not the per-token sticker. The win from this price war isn’t “spend the same and do 8× more” by accident — it’s choosing whether to bank the savings or reinvest them, on purpose. Cheap AI rewards discipline just as much as expensive AI did.

The bottom line

In a single day, the cost of running capable AI fell roughly 80%, and the labs signaled that the race is now about price as much as intelligence. For anyone actually shipping with AI, that’s a bigger story than any leaderboard crown — because it directly changes what’s profitable to build. Re-price your stack, route by job, keep an eye on total spend, and this price war becomes your tailwind.

Check the numbers yourself

Dig into current pricing:

Are you going to bank the savings or reinvest them into more output? And which model are you moving your heavy tasks to? Tell me in the comments.

Sources: OpenAI, xAI, and vendor pricing pages via BenchLM and AI Pricing Guru, July 2026. Prices are per million tokens (input / output) and change frequently — verify before you budget.

Leave a comment

Your email address will not be published. Required fields are marked *