What AI Coding Models Actually Cost Right Now

Bar chart of AI coding model output prices per million tokens: Claude Fable 5 at $50, Opus 5 at $25, GPT-5.6 Sol at $20 (promotional), Sonnet 5 at $10, Qwen3.8-Max at $6, Gemini 3.7 Flash at $3.75 (doubles in January), DeepSeek V4-Flash-Max at $0.20 — a 250x spread.

In ten days this month, OpenAI cut its flagship coding model by a third, Google launched a model at exactly half its predecessor’s price, and DeepSeek quadrupled its rates. Three moves, three directions, all in August. The story everyone reached for was “a price war is driving AI coding costs to zero,” which is roughly half right — and the half that’s wrong is the half that would cost you money. Here’s what everything actually costs as of today, where the real value sits, and which of these numbers has an expiry date attached.

The short version

  • Prices are not uniformly falling. DeepSeek went up about 4× on 16 August and introduced peak/off-peak rates.
  • Two of the cheapest headline numbers expire. Gemini 3.7 Flash doubles on 1 January 2027. GPT-5.6 Sol’s cut is promotional “at least through” 21 November 2026.
  • Anthropic made a discount permanent — Sonnet 5 stays at $2/$10 instead of rising to $3/$15 on 1 September, which complicates the “Anthropic is losing on price” story.
  • The FT story is about Claude Fable 5, not Opus 5. That distinction matters, and a lot of the commentary got it wrong.
  • The spread is the real finding: roughly 250× between the cheapest and dearest output prices, for under 30 points of benchmark difference.

All prices below are US dollars per million tokens, checked on 25 August 2026. Model pricing changes fast enough that you should treat any figure older than a few weeks as fiction — including this one, eventually.

What actually changed this month?

OpenAI cut GPT-5.6 Sol on 21 August, from $5/$30 to $4/$20. Worth being precise, because the widely repeated “more than 20%” is a blend: input fell exactly 20%, output fell 33%. If your workload is generation-heavy — which coding is — you got the better end of that. It’s also the third cut in the GPT-5.6 family inside a month, after Luna dropped 80% and Terra 20% on 30 July.

Google launched Gemini 3.7 Flash on 13 August at $0.75/$3.75 — exactly half Gemini 3.6 Flash’s $1.50/$7.50. The important detail is on Google’s own pricing page rather than in the press coverage: this is an introductory rate through 31 December 2026, and on 1 January 2027 it returns to $1.50/$7.50. Build your Q4 cost model on the current number and your New Year’s Day is a 100% increase.

DeepSeek went up, effective 16:00 UTC on 16 August, alongside the V4 lineup release. V4-Pro output moved from $0.87 to $3.96 at peak and $1.98 off-peak. That’s roughly a 4× increase at peak, and it came with a new peak/off-peak structure — 01:00–04:00 and 06:00–10:00 UTC on weekdays are peak; everything else is exactly half. DeepSeek framed it as enabling flexible workload scheduling. If your jobs are batch and can wait, that’s genuinely useful. If they’re interactive, it’s just a price rise.

Anthropic did the opposite of what you’d expect from the coverage. Sonnet 5 launched at $2/$10 as introductory pricing due to rise to $3/$15 on 1 September. That increase has been cancelled — $2/$10 is now the standard price. A permanent cut is a stronger signal than a temporary promotion, and it went largely unreported.

One correction while we’re here: reports that Moonshot has split Kimi K3 into separate general and coding subscription tiers are premature. The split is announced in the plan documentation, with existing subscribers keeping all benefits and no implementation date given. It hasn’t happened yet.

What does everything cost right now?

ModelInputOutputNotes
Claude Fable 5$10.00$50.00Anthropic’s top tier
Claude Mythos 5$10.00$50.00
Claude Opus 5$5.00$25.00
GPT-5.6 Sol$4.00$20.00Promo to 21 Nov 2026
Kimi K3$3.00$15.00$0.30 cached input
GPT-5.6 Terra$2.50$15.00
Claude Sonnet 5$2.00$10.00Now permanent
Qwen3.8-Max$2.00$6.00See caveat below
Z.ai GLM-5.3$1.40$4.40$0.26 cached input
DeepSeek V4-Pro (peak)$1.32$3.96$0.044 on cache hit
GPT-5.6 Luna$1.00$6.00
Gemini 3.7 Flash$0.75$3.75Doubles 1 Jan 2027
DeepSeek V4-Pro (off-peak)$0.66$1.98Half the peak rate
Gemini 3.5 Flash-Lite$0.30$2.50

Two honest caveats. Qwen3.8-Max is the weakest row here — Alibaba Cloud’s Model Studio pricing page wouldn’t load for me on repeated attempts, so $2/$6 comes from OpenRouter’s listing and secondary reporting rather than the vendor page. Verify it in the console before you build a budget on it. And if you see GPT-5.6 Sol listed at $2/$10 anywhere, that’s a routing-level discount or a rendering artefact — OpenAI’s own docs say $4/$20.

Is anyone actually abandoning the expensive models?

The Financial Times reported on 23 August that Anthropic’s best model is struggling to attract users while cheaper tools thrive. It’s a real story with real data behind it, and almost every summary I read got one detail wrong: it’s about Claude Fable 5, the $10/$50 flagship — not Opus 5, which sits at $5/$25 and appears to be gaining share. Several write-ups asserted the opposite of what the underlying numbers show.

The measured part comes from Ramp’s AI Index, built on corporate card data across roughly 70,000 companies. In July, Fable 5 accounted for about 6% of Anthropic tokens purchased but 11.4% of Anthropic model-attributed spending — the gap is just the price premium doing its work. By 23 August that had settled near 11%. That’s a plateau, not a collapse. Meanwhile Anthropic’s overall share of US businesses rose to 43.5%, up 1.1 points month on month, ahead of OpenAI’s 39.7%.

A second, independent dataset points the same way: Vercel’s AI Gateway saw Anthropic take 65.1% of gateway spending on 30% of tokens in July. And the single largest Anthropic line item in Ramp’s data wasn’t any flagship — it was Opus 4.8 at 28%, an older and cheaper model.

Both datasets have the same blind spot and it’s a big one: neither Ramp nor Vercel can see Anthropic’s direct enterprise contracts, which is arguably where a flagship model actually sells. Ramp’s own analysts also note the Fable sample skews more technical than their typical panel. The defensible version of this story is narrow: the expensive flagship is a small and flat share of card-and-gateway-visible spending. Not “nobody is buying it.”

The revenue figures in the FT piece — annualised revenue of $65bn in July, 6,000 customers above $100k a year — come from “people with knowledge of the matter,” not disclosed financials. Treat them accordingly. And the most quotable line in the coverage, that most people “don’t need to operate at the frontier,” is from an investor at a firm that has put around $1bn into Anthropic. That doesn’t make it wrong. It does make it a position rather than a finding.

What do you actually get for the extra money?

Here’s the SWE-bench Pro leaderboard as of 10 August, next to output price. Read it for shape, not for precise ordering — these are vendor-reported aggregate scores rather than a controlled harness, and on Scale’s standardised set with common scaffolding the ordering changes materially.

ModelSWE-bench ProOutput $/M
Claude Fable 580.0%$50
Claude Opus 4.869.2%$25
Qwen3.8-Max67.7%$6
GPT-5.6 Sol64.6%$20
Claude Sonnet 563.2%$10
GPT-5.6 Luna62.7%$6
GLM-5.262.1%$3
DeepSeek V4-Flash-Max52.6%$0.20

The number worth carrying away: the price spread across this set is roughly 250×. The capability spread is under 30 percentage points. Fable 5 buys you about 17 points over Sonnet 5 for five times the output price. Whether that’s worth it depends entirely on what those 17 points are doing — on a task where a wrong answer costs an hour of debugging, they’re cheap; on bulk refactoring where you review everything anyway, they’re not.

Be careful with points-per-dollar arithmetic, including mine. Dividing benchmark score by output price ignores input cost and cache economics entirely, which flatters models with cheap output and expensive input. It also treats a benchmark point as a linear good, which it isn’t.

How to actually choose

  • Price your own workload, not the sticker. Coding is output-heavy, so weight output price accordingly — that’s why Sol’s 33% output cut matters more than its 20% input cut.
  • Check the expiry date on any cheap number. Gemini 3.7 Flash doubles on 1 January 2027; Sol’s rate is promotional to 21 November 2026 with no stated price after.
  • Use cache pricing if your prompts repeat. DeepSeek’s cache-hit input is $0.044 against $1.32 on a miss — a 30× difference that dwarfs most model-choice decisions.
  • Consider off-peak scheduling. If your DeepSeek workload is batch, avoiding 01:00–04:00 and 06:00–10:00 UTC on weekdays halves the bill.
  • Don’t pay frontier prices for non-frontier tasks. The largest single Anthropic line item in the Ramp data is an older, cheaper model, which tells you something about what experienced buyers actually do.
  • Re-check quarterly. Four material price changes landed in one month. Anything you decided in June is already stale.

The useful mental shift is away from “which model is best” and toward “which model is cheapest for the tasks where cheapest is fine, and which is worth paying up for on the rest.” Almost nobody needs one model. Routing between two — a cheap one for bulk work and an expensive one for the parts that bite — beats picking a winner, and it’s the pattern the spending data suggests people are quietly converging on anyway. It’s the same instinct behind loading capability only when it’s needed rather than paying for it on every request.

Prices verified against vendor documentation on 25 August 2026. Benchmark figures from the SWE-bench Pro leaderboard dated 10 August 2026. Spending data from Ramp’s AI Index (July 2026) and Vercel AI Gateway (July 2026).

Leave a comment