
In late 2021, getting GPT-3-level answers out of an AI model cost around $60 for every million tokens processed.
By 2024, that same quality of output cost about six cents. Same intelligence, one-thousandth of the price. Then, sometime in early 2026, part of that curve started climbing back up.
What "Cost Per Intelligence" Actually Measures
Comparing AI prices model-by-model is misleading, because the models themselves keep changing. Cost per intelligence fixes the other variable: it asks what it costs, at any given moment, to buy a specific, fixed level of output quality, regardless of which model happens to deliver it.
Measured in dollars per million tokens, standardized against a benchmark quality level (like "GPT-3.5-equivalent")
Lets you compare 2021's best model to 2026's cheapest model on equal footing
Strips out marketing and focuses on what a fixed unit of capability actually costs to buy
The Number That Defined AI Economics
For three straight years, this number only moved one direction. Multiple independent analyses converge on the same shape, even if the exact multiples differ:
| Source | Finding |
|---|---|
| a16z, "LLMflation" (2024) | GPT-3-equivalent output fell from ~$60 to ~$0.06 per million tokens, 2021–2024 — roughly a 1,000x drop |
| Stanford HAI, AI Index 2025 | GPT-3.5-equivalent quality fell from $20 to $0.07 per million tokens in about 18 months — a 280-fold cut |
| Epoch AI, 2025 | Median benchmark price-performance decline of ~50x per year, accelerating to ~200x per year since January 2024 |
That's the trend line that made AI features commercially viable at consumer scale in the first place. It's also the trend line that, per newer 2026 data, no longer applies evenly across the market.
The Market Split In Two
Industry pricing trackers now describe 2026 as a market moving in two directions at once, not one. According to BenchLM's Token Price Index, mid-tier and budget models are still falling fast, down roughly 36% year-over-year as of mid-2026. The frontier tier, meanwhile, has done the opposite.
Axis Intelligence Research's LLMflation Index, which tracks the newest top-capability model's price relative to GPT-4's March 2023 launch, peaked above 1,000 in mid-2025 (frontier inference over 1,000x cheaper than GPT-4 at launch). By July 2026 it had fallen back to around 333, as newer flagship models replaced cheaper ones at higher price points, still far cheaper than 2023, but roughly three times pricier than the market's own 2025 low point.
What This Means For Your AI Bill
The practical takeaway isn't "AI got more expensive" or "AI got cheaper," it's that those two things are now true simultaneously, depending on which tier you're buying:
Routine, high-volume tasks belong on budget or mid-tier models, where deflation is still doing the work for you
Frontier models are worth paying the rising premium for only when the task genuinely needs the extra capability
The lowest per-token price isn't always the lowest total cost, a cheaper model that needs more retries or corrections can cost more per finished task
Cost Per Intelligence: FAQ
It's a way of pricing AI not by the model but by a fixed level of output quality — for example, what it costs today to get GPT-3.5-level answers, versus what it cost when that quality was the best available. Because model quality keeps improving, tracking price per model is misleading; tracking price per quality level shows the real trend.
For most use cases, yes, dramatically. Andreessen Horowitz's LLMflation analysis found GPT-3-equivalent output fell from about $60 to $0.06 per million tokens between 2021 and 2024, a roughly 1,000x drop. Stanford HAI's AI Index 2025 separately measured a 280-fold drop for GPT-3.5-equivalent quality in about 18 months. Budget and mid-tier models are still falling in price through 2026.
Because two different things are happening at once. Budget and mid-tier model pricing keeps falling, roughly 36% year-over-year according to BenchLM's Token Price Index. But frontier-tier pricing, the newest top-capability model at any given moment, has risen since January 2026 as each generation adds capability and commands a premium. Reasoning models also generate many more internal tokens per answer than they used to, which can raise a total bill even when the per-token price is falling.
It's a tracker, popularized by a16z and continued by outlets like Axis Intelligence Research, that measures the blended per-token price of the frontier-tier model relative to GPT-4's March 2023 launch price. It peaked above 1,000 (meaning frontier inference was over 1,000x cheaper than GPT-4 at launch) in mid-2025, then fell back toward 333 by July 2026 as newer, pricier flagship models replaced cheaper ones.
Route routine, high-volume work to budget or mid-tier models, where prices are still falling fast, and reserve frontier-tier models for tasks that genuinely need the extra capability. The cheapest per-token price isn't always the cheapest finished task, since a weaker model that needs retries or corrections can cost more overall than a pricier model that gets it right the first time.
Jans Bock-Schroeder
Publisher & Founder of AI Angst
Coming from the world of art, photography, and the luxury market, Jans launched AI Angst in 2025 to explore the cultural, ethical, and psychological impacts of artificial intelligence. His work bridges creative vision with critical technology analysis, offering clarity in an era of rapid technological change.
Sources and Citations
This article is based on the following sources:
-
Andreessen Horowitz: "Welcome to LLMflation — LLM inference cost is going down fast" (2024)
Original source for the ~1,000x, 2021–2024 GPT-3-equivalent price decline.
https://a16z.com/llmflation-llm-inference-cost/ -
Stanford HAI: "Artificial Intelligence Index Report 2025"
Source for the 280-fold, 18-month price decline for GPT-3.5-equivalent quality.
https://hai.stanford.edu/ai-index/2025-ai-index-report -
Epoch AI: "LLM inference prices have fallen rapidly but unequally across tasks" (2025)
Source for the median 50x-per-year, accelerating to 200x-per-year, benchmark price decline.
https://epoch.ai/data-insights/llm-inference-price-trends -
Axis Intelligence Research: "AI Inference Cost Statistics 2026: The Market That Split in Two"
Source for the LLMflation Index's mid-2025 peak, its July 2026 reading of 333, and the frontier-versus-budget price divergence.
https://axis-intelligence.com/ai-inference-cost-statistics/ -
VoxBooster: "AI Inference Cost Statistics (2026): 50+ Data Points on the Price Collapse"
Consolidated cross-check of the a16z, Stanford HAI, and Epoch AI figures cited above.
https://voxbooster.com/blog/ai-inference-cost-statistics-2026/
Published: August 30, 2026. Sources verified at time of publication. All external links open in a new tab.


