TEHNOLOGIE

AI token prices have not fallen since April — and the frontier got 72% more expensive

olivLaw Psychohistory
Four NVIDIA H100 datacenter accelerator cards laid out in a row
NVIDIA H100 accelerator cards. Renting them on a one-year contract got roughly 40% more expensive between October 2025 and March 2026, and the memory shortage is the reason. Photo: Wikimedia Commons.

1. Why we re-ran the check

Since the end of April 2026, our /ai-hub page has displayed a table with the public token price of fourteen language models, plus a derived column — "intelligence per dollar" — dividing a leaderboard score by the blended cost. The table is labelled as of 2026-04 and is written by hand in code, not pulled automatically from vendors.

A hand-written table that does not update itself can be wrong in two ways: it can be wrong from the start, and it can fall behind. We checked both, separately, because they answer different questions. The first is a mistake of ours. The second is a measurement about the market — and, to our surprise, it is the interesting part.

A token is the chunk of text a model processes: roughly three quarters of a word in English, somewhat less in Romanian. The prices below are in dollars per million tokens, as each vendor publishes them. Where a single figure appears, it is the blended price: 70% input, 30% output, the usual proportion for a dialogue with long questions and short answers. It is the same convention the page uses.

2. First finding: the April table does not check out

Of fourteen rows, three name models that have never existed under those names:

  • "Llama 4 405B". The Llama 4 family published by Meta has three members — Scout, Maverick and Behemoth. There is no 405-billion-parameter model in it; 405B is the top model of the previous generation, Llama 3.1. The name on the page is a crossing of two generations.
  • "DeepSeek V3.5". The series went V3.1 → V3.2 (1 December 2025) → V4 (24 April 2026). There was no V3.5. The price shown, $0.27 / $1.10, is in fact exactly V3.1's price.
  • "Kimi K2 Pro". Moonshot has published K2, K2.5, K2.6, K2.7 Code and K3. Not a "K2 Pro".

Of the eleven remaining rows, which name real models, eight carried the wrong price. Only three were correct.

ModelShown on the page (in / out)Real price, April 2026Error
Claude Opus 4.7$15 / $75$5 / $253 times too high
Claude Sonnet 4.6$3 / $15$3 / $15correct
Claude Haiku 4.5$0.80 / $4$1 / $520% too low
GPT-5$5 / $15$1.25 / $104 times too high on input
GPT-5 Mini$0.50 / $2$0.25 / $22 times too high on input
GPT-4o$2.50 / $10$2.50 / $10correct
Gemini 2.5 Pro$3.50 / $10.50$1.25 / $102.8 times too high on input
Gemini 2.5 Flash$0.30 / $2.50$0.30 / $2.50correct
Grok 4$5 / $15$3 / $151.67 times too high on input
Qwen3-235B-A22B$0.50 / $1.50$0.455 / $0.901.67 times too high on output
Mistral Medium 3$2.70 / $8.10$0.40 / $26.75 times too high on input

The average blended price across the eleven real models: $6.795 on the page versus $3.828 in reality. The table inflated the cost of tokens by 77.5%.

The most instructive error is the first one. The $15 / $75 price was real — but for Claude Opus 4.1, until 24 November 2025, when Anthropic cut the frontier model's rate by 67%, to $5 / $25. Our table, written on 29 April 2026, reproduced a price that had been withdrawn five months earlier and attributed it to a model launched thirteen days before. It is not a figure invented out of thin air: it is a figure that used to be true and that nobody re-checked. The rest of the errors do not even have that excuse — $2.70 was never the price of Mistral Medium 3, which has cost $0.40 since launch.

The practical consequence is not cosmetic. The "intelligence per dollar" column divides a score by the blended price. With a denominator wrong by up to 6.75 times, the ordering in that column — the one thing a reader would use to pick a model — was arbitrary.

3. Second finding: since April, not one price has moved

Here we expected the boring part of the article: re-laying the table with today's prices and computing how far they had fallen in five months. They did not fall. Nor did they rise.

ModelApril 20263 September 2026Change
Claude Opus 4.7$5 / $25$5 / $250%
Claude Sonnet 4.6$3 / $15$3 / $150%
Claude Haiku 4.5$1 / $5$1 / $50%
GPT-5$1.25 / $10$1.25 / $100%
GPT-5 Mini$0.25 / $2$0.25 / $20%
GPT-4o$2.50 / $10$2.50 / $100%
Gemini 2.5 Pro$1.25 / $10$1.25 / $100%
Gemini 2.5 Flash$0.30 / $2.50$0.30 / $2.500%
Grok 4$3 / $15withdrawn from the cataloguepulled on 15 May 2026
Qwen3-235B-A22B$0.455 / $0.90$0.455 / $0.900%
Mistral Medium 3$0.40 / $2$0.40 / $20%

Ten out of ten — of the eleven rows the panel was publishing in April that named real models. No change, in any position, in five months. The eleventh, Grok 4, neither rose nor fell: it was pulled entirely on 15 May 2026. The old name does not return an error; it routes requests to Grok 4.3, at $1.25 / $2.50 — a blended cost of $1.63 instead of $6.60. It is the only movement in the table, and it happens to go the buyer's way.

Some of the unchanged prices have been frozen for far longer than that: Gemini 2.5 Flash-Lite has cost $0.10 / $0.40 since its launch on 17 June 2025, and GPT-5 nano has cost $0.05 / $0.40 since 7 August 2025. Over thirteen months without a single change, in a market that the trade press was describing, a few months earlier, as getting cheaper by tens of times a year.

4. What actually moved: the menu — and upwards

Prices did not change, but the catalogue did. And the new models did not arrive below the price of the ones they replace, as they used to until last year. They arrived above.

VendorTop of the range in AprilTop of the range in SeptemberBlended costChange
OpenAIGPT-5 ($1.25 / $10)GPT-5.6 Sol ($4 / $20, promotional rate)3.875 → $8.80+127%
GoogleGemini 2.5 Flash ($0.30 / $2.50)Gemini 3.6 Flash ($0.75 / $3.75)0.96 → $1.65+72%
AnthropicOpus 4.7 ($5 / $25)Fable 5.1, a new tier above Opus ($10 / $50)11.00 → $22.00+100%
GoogleGemini 2.5 Pro ($1.25 / $10)Gemini 3.1 Pro ($2 / $12)3.875 → $5.00+29%
xAIGrok 4 ($3 / $15)Grok 4.6 ($2 / $6)6.60 → $3.20−52%

The median: +72% in five months. Two of the figures in that table are more fragile than they look, however, and the fragility pushes the same way. GPT-5.6 Sol's rate, $4 / $20, is flagged by OpenAI as promotional and guaranteed only until 21 November 2026; the standing rate for the top tier, inherited from GPT-5.5, is $5 / $30 — that is $12.50 blended and +223% against April. And $0.75 / $3.75 for Gemini 3.6 Flash is the rate valid until 31 December 2026: for 1 January 2027, Google has already published double, $1.50 / $7.50, which would mean +244%. With the standing rates in place of the promotional ones, the median is not +72% but +100%. We used the lower figure throughout, the one that is actually paid today.

The exact dates matter, because they show this is not a calendar coincidence:

  • 23 April 2026 — OpenAI launches GPT-5.5 at $5 / $30. The top model goes from $1.25 to $5 on input: four times.
  • 15 May 2026 — xAI withdraws Grok 4 and routes the old name to Grok 4.3, at $1.25 / $2.50. The first of the two clear counter-examples in the window.
  • 9 June 2026 — Anthropic opens a tier above Opus, at $10 / $50, double Opus on both ends. Opus's own rate stays at $5 / $25, unchanged since the November 2025 cut, across four successive generations.
  • 9 July 2026 — GPT-5.6 replaces GPT-5.5. The top tier, Sol, is listed at $4 / $20: below its predecessor, but with an explicit note that this is a time-limited promotion.
  • 30 July 2026 — OpenAI cuts two of the GPT-5.6 tiers. The cut applies against its own 9 July launch prices, not against April.
  • 11 August 2026 — Anthropic cancels an increase it had already announced: Sonnet 5 was due to go from $2 / $10 to $3 / $15 on 1 September, and instead the introductory rate becomes permanent. The second counter-example, and the hardest one to square with the shortage explanation.
  • 12 August 2026 — xAI launches Grok 4.6 cheaper than Grok 4.
  • 16 August 2026 — DeepSeek raises its rates. It is the largest increase in the window, at the cheapest vendor in the group.

The last point deserves detail, because it is the opposite of what anyone expected from DeepSeek. The vendor gave notice on 6 August, published the new grid on 13 August and applied it on 16 August, at 16:00 UTC. On V4-Flash, the entry model, the rate for text not already cached goes from $0.14 to $0.22 off-peak and to $0.44 at peak; output goes from $0.28 to $0.66 and $1.32 respectively.

DeepSeek V4-FlashUntil 16 AugustOff-peakAt peak hours
Input$0.14$0.22 (+57%)$0.44 (+214%)
Output$0.28$0.66 (+136%)$1.32 (+371%)
Blended cost$0.182$0.352 (+93%)$0.704 (+287%)

V4-Pro, the large model, has the same structure: $0.66 / $1.98 off-peak and $1.32 / $3.96 at peak. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday.

There are two new things here, not one. The first is the size of the increase, at the vendor that defined the cheap end of the market. The second is its shape: a rate that depends on the hour at which you send the request is how electricity or transport capacity is sold, not software. A vendor gets there when the binding constraint is no longer what the model costs to develop, but how much capacity is free at a given moment of the day.

5. The rate of change, in three measures

"How fast did prices move" has no single answer, because it depends on what you hold fixed. Three different questions, three different figures, all over the window 23 April – 3 September 2026:

The questionOver 5 monthsEquivalent annual rate
What does the exact model I was using in April cost today?0.0%0.0%
What does the cheapest listed model cost?0.0%0.0%
What does the best thing on the market cost? (median of 5 vendors)+72%does not annualise

The last figure needs a caveat, otherwise it misleads. A 72% rise over five months would annualise mathematically to +267%, but that number would mean nothing: it extrapolates a sample of five vendors over a single window, and frontier rate jumps are rare events, not a continuous rate. The honest figure is the five-month one. We left it un-annualised on purpose.

The second row has an exception worth stating too. The cheapest listed rate in the tracked group — $0.05 / $0.40 on GPT-5 nano — genuinely did not move at all. But the cheap end of the market, taken as a whole, did move: on 16 August, DeepSeek raised V4-Flash by 93% off-peak and by 287% at peak. Anyone who had picked the cheapest vendor, rather than the cheapest model on a fixed list, felt this window.

The benchmark against which all three should be read: market analyses for 2021–2025 measured, at equal capability, a median decline on the order of tens of times per year, with a very wide range across tasks. However exactly that range was measured, the order of magnitude was "getting cheaper fast". Over the five months analysed here, the order of magnitude is zero — and at the frontier, negative for the buyer.

6. The causal chain

The cause. The high-speed memory stacked next to the graphics processor — the component that holds the model while it runs — is made by a very small number of suppliers, and demand has outrun what they can ship. That cost was passed on: the one-year contract rent for H100-class accelerators rose from roughly $1.70 to roughly $2.35 per hour between October 2025 and March 2026, close to 40%. This is the price of renting, not of buying — model vendors generally rent the machines they run on.

The mechanism. Rent got more expensive at the next level up as well. Amazon raised the price of reserved GPU capacity blocks by 15% on 4 January 2026 and by a further 20% or so on 1 July 2026 — the mechanism by which a customer secures machines in advance for training or for traffic peaks. These are the first increases of their kind at a major infrastructure vendor in almost two decades; until now the direction had been one-way. The distinction matters, because not everything GPU-related got more expensive there: on-demand rates for the same machines had been cut by up to 45% in June 2025. The direction reversed exactly on capacity that is reserved ahead of time, which is exactly where a shortage is felt. A rising input cost lifts the floor below which a list price cannot go.

The effect. When unit cost stops falling, a vendor has two options. It can hold the rate and leave the old model in place — that explains the ten zeroes in the table in section 3. Or it can bring the extra capacity to market at a higher price, under a new name — that explains GPT-5.5 at four times the price of GPT-5, and a new tier above Opus. Both behaviours show up simultaneously, at the same vendors, in the same window. That is the signature of a supply constraint, not of an isolated commercial decision.

There is also a third option, the most brutal: raising the price of a model already in the catalogue. A single vendor took it in this window, DeepSeek, and that is precisely why it matters more than its market size — it was the only one that paid the reputational cost of a direct increase, instead of hiding it inside a new generation.

7. The counter-hypothesis

The explanation above may be wrong, and there is a serious alternative: the cheapening has not stopped, it has moved off the price list.

The argument is solid, but its figures need precision. The discount for text already submitted once — prompt caching — goes up to 90% of the input rate at Anthropic, 75% at Google and 50% at OpenAI, where 90% applies only to recent tiers. So it is not a general 90%, as it is often quoted. Batch processing, without an immediate answer, halves the rate at all three. And open-weight models, taken from a third-party host rather than from their author, cost 50–90% less — but that is a permanent structural gap between two markets, not a cheapening that happened inside the window analysed. On this reading, the effective price paid by a customer who optimises their requests has kept falling, and the list rate has become a shop window nobody updates any more because nobody pays exactly that.

The counter-hypothesis also has two hard facts on its side, both from the window analysed. On 15 May, xAI withdrew Grok 4 and routed the old name to a model 75% cheaper on blended cost. On 11 August, Anthropic cancelled an increase it had already communicated — Sonnet 5 moving from $2 / $10 to $3 / $15, scheduled for 1 September — and made the introductory rate permanent. A vendor squeezed by a real shortage does not walk back an increase it has already announced. At Anthropic, in August at least, cost pressure was not enough to pass into price.

The two explanations are not mutually exclusive and both are probably partly true. What separates them is an observation anyone can make: if real cheapening were continuing unimpeded, a vendor would have every interest in displaying it, because the list rate is the main instrument of comparison between competitors. The fact that not one of the ten models still listed moved anything in five months, in a market where displaying a cut costs nothing, is hard to explain by update laziness alone. The figure that would settle the dispute — the average effective cost paid per million tokens, market-wide — is published by nobody.

8. What this means in practice

  • Cutting costs by waiting no longer works. The strategy "write the application on the expensive model, the price will drop by launch anyway" rested on a rate that, over the last five months, has been zero. Budget at today's price.
  • The old model is the cheap option, and it stays. A model that has not gone up in five months and that the vendor keeps listed is the most predictable cost line available. The risk is not the price but the withdrawal: on 11 June 2026, OpenAI announced the removal from the API, on 11 December 2026, of the dated snapshots in the GPT-5 family — those pinned to 7 August 2025, plus the Pro variant from 6 October 2025 — replaced by GPT-5.6. A stable rate does not help if the model disappears.
  • The hour you send the request has become a cost variable. At DeepSeek, the same request costs twice as much between 01:00 and 04:00 and between 06:00 and 10:00 UTC than at other times. Workloads that can wait — overnight summaries, reprocessing, indexing — are worth scheduling outside those windows. For now it is a single vendor, but it is the first to have moved the rule.
  • Jumping to the new generation is now a cost decision, not a free one. Until 2025, the new version came out better and cheaper. Since April 2026, at three of the four vendors that refreshed their top of the range, it comes out better and more expensive — by up to 127% on blended cost, or 223% if OpenAI's promotional rate expires in November. It is worth proving the workload actually needs it.
  • A hand-written price table breaks silently. Ours sat wrong for four months without signalling anything. A table that updates itself gets things wrong too, but it gets them wrong loudly.

9. Falsifiable predictions

Each prediction carries an explicit horizon, a probability in calibrated language and a public source at which anyone can check it. They are written so they can be contradicted.

PredictionHorizonProbabilityHow to verify
F1. Claude Opus stays at $5 / $25 per million tokens31 December 2026very likely (0.82)Anthropic's public pricing page
F2. At least 8 of the 10 models still listed carry the same rate as today31 December 2026likely (0.70)The public pricing pages of the 6 vendors
F3. No major vendor cuts the list rate of an existing model by more than 30%31 December 2026roughly even odds, leaning likely (0.60)Public pricing announcements
F4. OpenAI's next frontier model has an input rate of at least $4 per million tokens31 March 2027likely (0.68)OpenAI's public pricing page
F5. GPT-5.6 Sol's rate exceeds $4 on input after the promotion expires, on 21 November 202631 December 2026roughly even odds, leaning likely (0.55)OpenAI's public pricing page
F6. A second vendor, after DeepSeek, introduces peak-hour differentiated pricing30 June 2027unlikely (0.35)Public pricing pages
F7. The dated snapshots in the GPT-5 family are effectively removed from the API on the announced date31 December 2026very likely (0.85)OpenAI's public deprecation list
F8. One-year contract rent for datacenter accelerators does not fall back below its October 2025 level30 June 2027likely (0.72)Public component market reports
F9. No major vendor publishes the average effective cost paid per million tokens31 December 2026very likely (0.88)Vendors' public reports and documentation
F10. A model launched after 1 September 2026 comes in below $0.05 per million input tokens at a major vendor30 June 2027unlikely (0.38)Public pricing pages

10. What would invalidate this analysis

  • The sample. Eleven models, chosen because they were on our April page, not because they represent the market. US vendors are over-represented and open models hosted by third parties — where price competition is fiercest — are almost entirely missing.
  • The window. Five months. A single round of frontier increases can be a calendar coincidence between a few launches rather than a regime change. The same measurement in six months is the test.
  • The list rate is not the invoice. We measure what vendors display, not what customers pay. Caching discounts, batches and negotiated contracts can shift the real cost by an order of magnitude, and those figures are not public. This is the counter-hypothesis in section 7, and it remains unsettled.
  • Promotional rates poison any series. Two of the five figures in section 4 are time-limited prices: OpenAI's promotion expires on 21 November 2026, and the Gemini Flash rate doubles on 1 January 2027. We chose the price paid today everywhere, which is the lower one — but the same measurement made in February, with exactly the same method, will give a noticeably higher result without anything in the market having changed in between.
  • Comparing across generations is imperfect. When we say the frontier got 72% more expensive, we are comparing models with different capabilities. A model twice as expensive that solves three times as many tasks is not a price increase for the buyer. We do not have an independent capability measure we trust enough to compute that ratio — and, as section 2 shows, a bad ranking is worse than none.

11. Sources

Rates and announcements: Anthropic's public pricing, the announcement of the 67% Opus cut, November 2025, the announcement of the new tier above Opus; the announcement making Sonnet 5's introductory rate permanent; the GPT-5.5 announcement, OpenAI's pricing grid (where the promotion's end date is flagged) and OpenAI's deprecation list; Gemini pricing, with the dates on which it changes; the xAI model list, where the withdrawal of Grok 4 is recorded; DeepSeek pricing and the announcement of the 16 August 2026 increase; Mistral pricing; Alibaba Model Studio pricing. Infrastructure cost: Amazon's pricing for GPU capacity blocks.

Verification of model names: the Llama 4 family as announced by Meta (Scout, Maverick, Behemoth — no 405B). The historical rate of cheapening: Epoch AI's series on inference price trends. Regulatory framework: Regulation (EU) 2024/1689.

The audited table is the one on /ai-hub, in the form published between 29 April and 3 September 2026. It was corrected alongside the publication of this analysis and now shows, next to the current price, the April price and the difference — so that the error stays visible rather than erased.

Disclaimer: This material is for information and analysis. It does not constitute investment advice, a technology procurement recommendation or a legal assessment. The prices quoted are list rates published by vendors and verified on 3 September 2026; they change frequently and may differ significantly from the effective billed cost, which depends on discounts, contracts and usage patterns. The probabilities are explicit estimates, written so that events can contradict them, not certainties.