Scale · founder · 7 min read
Three Frontier Models in 48 Hours, All Cheaper. Here's the Part That Reaches Your Bill.
Grok 4.7, Claude Opus 5.5 and GPT-6 Sol/Luna landed in two days at lower prices. The price war is one tier below the flagships — the tier your tools run on.
On September 21, xAI shipped Grok 4.7. On September 22, Anthropic released Claude Opus 5.5, and roughly an hour later OpenAI released GPT-6 Sol and GPT-6 Luna. Three vendors, two days, and every one of them either cut prices or held them flat while shipping a better model.
You don’t buy tokens. You buy a $25/month Lovable seat or a $20 Cursor subscription. So the useful question isn’t which model won a benchmark. It’s which of these changes actually reaches your invoice, and when.
What the price table looks like now
Per million tokens, as of September 22:
| Model | Input | Cached input | Output |
|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.01 | $0.50 |
| Grok 4.7 | $2 | $0.50 | $6 |
| GPT-6 Sol | $2 | $0.20 | $10 |
| Claude Opus 5.5 | $4 | $0.20 | $20 |
| GPT-6 Astra | $10 | $1 | $50 |
| Claude Fable 5.1 | $10 | $0.25 | $50 |
Three things to notice.
GPT-6 Sol and Luna are roughly half the price of their GPT-5.6 equivalents. GPT-5.6 Sol was $4/$20; GPT-6 Sol is $2/$10. GPT-5.6 Luna was $0.20/$1.20; GPT-6 Luna is $0.10/$0.50. And GPT-5.6 has a scheduled 25% increase in November, so the real gap against what you’d pay in six weeks is wider than the table shows.
Opus 5.5 cut 20% off a price that hadn’t moved in five releases. Opus 4.5 through Opus 5 all sat at $5/$25. Opus 5.5 is $4/$20. The cache read price fell 60%, from $0.50 to $0.20 — and in long agentic sessions, most of your input tokens are cache reads, so that line matters more than the headline.
Grok 4.7 held at $2/$6 while getting meaningfully better at coding. xAI’s own numbers: CursorBench 4.0 went from 40.4% to 46.3%, DeepSWE v1.1 from 65.2% to 71.0%. Vendor benchmarks on vendor evals, treat accordingly — but the price didn’t move, which is the part that isn’t a claim.
The war is happening one tier below the flagships
Look at the bottom two rows. GPT-6 Astra and Claude Fable 5.1 are both still $10/$50. Nobody cut the top tier.
Everything moved in the band underneath: $2 to $4 input, $6 to $20 output. That band is not an accident. It is where agentic coding actually runs — long sessions, lots of file reads, lots of tool calls, thousands of tokens per step. It is also where the tools you pay a flat monthly fee to do most of their work, because the flagship tier is too expensive to put behind a $25 subscription.
So the vendors are competing hardest on exactly the layer that sits underneath your bill. That’s good news for you, with a delay attached.
What actually reaches you, and when
Almost nothing, this week.
Your Lovable seat is still $25. Your Cursor subscription is still $20. Model price cuts reach a flat-fee product as margin first, and as product second. What you should expect over the next one to three months, in rough order of likelihood:
- More generous limits at the same price. The cheapest way for a tool to pass a cut along is to stop throttling you. This is also the least visible — you notice you stopped hitting the wall, not that anything changed.
- A better default model on the cheap plan. If your tool’s free or entry tier was quietly running something weak, GPT-6 Luna at $0.10/$0.50 makes a real model affordable there.
- A new cheaper tier, aimed at people who bounced off the $25 entry price.
- Nothing at all, and the vendor keeps the margin. This is common and not dishonest — most of these companies are not profitable.
The one place you see it immediately is if you pay per token. If you’re on API billing for Claude Code or driving anything through OpenRouter, your next invoice is genuinely 20-50% smaller for the same work. That’s it. That’s the whole list of people who get an instant win.
Both major agents also switched defaults on day one: Opus 5.5 became the default Opus model in Claude Code 2.1.280, and GPT-6 Sol became Codex’s default starting model. If you use either, you got a different model under you on Tuesday without doing anything. That’s worth knowing before you conclude your tool “got weird this week.”
The honest caveat nobody in the press release will mention
Simon Willison ran Opus 5.5 at its “max” thinking level and it failed to return an answer at all — twice. It reasoned so long about a simple SVG drawing that it hit the 128,000-token output ceiling mid-thought and stopped. Each failed attempt cost $2.56 and took close to 20 minutes.
That is one test, on one prompt, from one person. But it points at something real: effort levels are now a pricing decision as much as a quality one, and the highest setting is not automatically the best one. If your tool exposes a thinking-effort or reasoning-level control, the top notch is a thing to test on your own work, not a thing to leave on.
What to do this week
Nothing urgent. No action here is time-sensitive, and that’s worth saying plainly because three launches in 48 hours creates a feeling of needing to move.
If you pay per token, check your default model. Anything still pointed at GPT-5.6 Sol is paying double for a model that a cheaper one now matches. Same logic for GPT-5.6 Terra, which is now priced identically to GPT-6 Sol with no remaining reason to prefer it.
Don’t hardcode a version string. GPT-5.5 retires from Codex on October 14. GPT-5.4 was pulled at the end of August. Treat six months as the shelf life of any model name in a script or a scheduled task, and read what happens when a plan shrinks without a price change for why that keeps happening.
Watch your tool’s changelog, not the model vendors’. The thing that changes your experience is the day Lovable or Cursor or Replit swaps what’s running underneath. That’s covered in which model powers your vibe coding tool, and it moves more often than any of these companies announce.
The limit of this
A price war in the mid tier does not mean AI coding is getting cheaper for you. The two things that determine your actual bill are how many tokens your work consumes and what your vendor charges for a seat, and both have been going up all year. Cheaper tokens have historically been absorbed by longer agent runs rather than passed along — the same pattern as the Fable 5.1 cache price cut three weeks ago, which almost nobody noticed on their invoice.
What today actually buys you is optionality. The tier below the flagships is now good enough and cheap enough that reaching for the $50-output model is a deliberate choice instead of a default. For most of what a non-technical founder builds, it was never the right default anyway.
Related guides
founder · 7 min read
Anthropic's June 15 Billing Change: What It Means If You Build With Claude
On June 15, Anthropic splits interactive and programmatic usage into separate pools. Here's what changes, and what doesn't, if you don't write code.
founder · 8 min read
NewClaude Fable 5.1: The Price Cut Hiding Behind an Unchanged Price Tag
Anthropic kept Fable's sticker price identical and cut the real cost 25-45%. Why your tool bill may move before your tool announces anything.
founder · 7 min read
Claude Fable 5 Is Here: What It Means for Vibe Coders
Anthropic's new top model lands in Claude Code and GitHub Copilot, free until June 22. What changes for the tools you build with, and what doesn't.
Enjoying this guide?
Get weekly practical guides, plus tool updates and implementation playbooks.