The Cost Of AI Intelligence Fell 1,000x In Three Years. In 2026, Part Of That Trend Quietly Reversed.

August 30, 2026 AI Angst avatar — a robot head with a distressed expression. JBS

A minimalist line chart glowing on a dark background, sloping steeply downward and then curving slightly back upward near the right edge, styled like a stock ticker display.AI Label

In late 2021, getting GPT-3-level answers out of an AI model cost around $60 for every million tokens processed.

By 2024, that same quality of output cost about six cents. Same intelligence, one-thousandth of the price. Then, sometime in early 2026, part of that curve started climbing back up.


What "Cost Per Intelligence" Actually Measures

Comparing AI prices model-by-model is misleading, because the models themselves keep changing. Cost per intelligence fixes the other variable: it asks what it costs, at any given moment, to buy a specific, fixed level of output quality, regardless of which model happens to deliver it.

  • Measured in dollars per million tokens, standardized against a benchmark quality level (like "GPT-3.5-equivalent")

  • Lets you compare 2021's best model to 2026's cheapest model on equal footing

  • Strips out marketing and focuses on what a fixed unit of capability actually costs to buy


The Number That Defined AI Economics

For three straight years, this number only moved one direction. Multiple independent analyses converge on the same shape, even if the exact multiples differ:

Source Finding
a16z, "LLMflation" (2024) GPT-3-equivalent output fell from ~$60 to ~$0.06 per million tokens, 2021–2024 — roughly a 1,000x drop
Stanford HAI, AI Index 2025 GPT-3.5-equivalent quality fell from $20 to $0.07 per million tokens in about 18 months — a 280-fold cut
Epoch AI, 2025 Median benchmark price-performance decline of ~50x per year, accelerating to ~200x per year since January 2024

That's the trend line that made AI features commercially viable at consumer scale in the first place. It's also the trend line that, per newer 2026 data, no longer applies evenly across the market.


The Market Split In Two

Industry pricing trackers now describe 2026 as a market moving in two directions at once, not one. According to BenchLM's Token Price Index, mid-tier and budget models are still falling fast, down roughly 36% year-over-year as of mid-2026. The frontier tier, meanwhile, has done the opposite.

Axis Intelligence Research's LLMflation Index, which tracks the newest top-capability model's price relative to GPT-4's March 2023 launch, peaked above 1,000 in mid-2025 (frontier inference over 1,000x cheaper than GPT-4 at launch). By July 2026 it had fallen back to around 333, as newer flagship models replaced cheaper ones at higher price points, still far cheaper than 2023, but roughly three times pricier than the market's own 2025 low point.

The floor and the ceiling are no longer moving together. Budget inference keeps getting closer to free, which is exactly what let AI features spread into products that couldn't have afforded them two years ago. But the newest, most capable model at any given moment now carries a real premium again, and reasoning models can burn far more tokens internally than what they show you, which quietly inflates a bill even when the sticker price per token looks unchanged.

What This Means For Your AI Bill

The practical takeaway isn't "AI got more expensive" or "AI got cheaper," it's that those two things are now true simultaneously, depending on which tier you're buying:

  • Routine, high-volume tasks belong on budget or mid-tier models, where deflation is still doing the work for you

  • Frontier models are worth paying the rising premium for only when the task genuinely needs the extra capability

  • The lowest per-token price isn't always the lowest total cost, a cheaper model that needs more retries or corrections can cost more per finished task


Cost Per Intelligence: FAQ

It's a way of pricing AI not by the model but by a fixed level of output quality — for example, what it costs today to get GPT-3.5-level answers, versus what it cost when that quality was the best available. Because model quality keeps improving, tracking price per model is misleading; tracking price per quality level shows the real trend.

For most use cases, yes, dramatically. Andreessen Horowitz's LLMflation analysis found GPT-3-equivalent output fell from about $60 to $0.06 per million tokens between 2021 and 2024, a roughly 1,000x drop. Stanford HAI's AI Index 2025 separately measured a 280-fold drop for GPT-3.5-equivalent quality in about 18 months. Budget and mid-tier models are still falling in price through 2026.

Because two different things are happening at once. Budget and mid-tier model pricing keeps falling, roughly 36% year-over-year according to BenchLM's Token Price Index. But frontier-tier pricing, the newest top-capability model at any given moment, has risen since January 2026 as each generation adds capability and commands a premium. Reasoning models also generate many more internal tokens per answer than they used to, which can raise a total bill even when the per-token price is falling.

It's a tracker, popularized by a16z and continued by outlets like Axis Intelligence Research, that measures the blended per-token price of the frontier-tier model relative to GPT-4's March 2023 launch price. It peaked above 1,000 (meaning frontier inference was over 1,000x cheaper than GPT-4 at launch) in mid-2025, then fell back toward 333 by July 2026 as newer, pricier flagship models replaced cheaper ones.

Route routine, high-volume work to budget or mid-tier models, where prices are still falling fast, and reserve frontier-tier models for tasks that genuinely need the extra capability. The cheapest per-token price isn't always the cheapest finished task, since a weaker model that needs retries or corrections can cost more overall than a pricier model that gets it right the first time.


Jans Bock-Schroeder, AI Expert and Founder of AI Angst

Jans Bock-Schroeder

Publisher & Founder of AI Angst

Coming from the world of art, photography, and the luxury market, Jans launched AI Angst in 2025 to explore the cultural, ethical, and psychological impacts of artificial intelligence. His work bridges creative vision with critical technology analysis, offering clarity in an era of rapid technological change.


Sources and Citations

This article is based on the following sources:

  1. Andreessen Horowitz: "Welcome to LLMflation — LLM inference cost is going down fast" (2024)
    Original source for the ~1,000x, 2021–2024 GPT-3-equivalent price decline.
    https://a16z.com/llmflation-llm-inference-cost/
  2. Stanford HAI: "Artificial Intelligence Index Report 2025"
    Source for the 280-fold, 18-month price decline for GPT-3.5-equivalent quality.
    https://hai.stanford.edu/ai-index/2025-ai-index-report
  3. Epoch AI: "LLM inference prices have fallen rapidly but unequally across tasks" (2025)
    Source for the median 50x-per-year, accelerating to 200x-per-year, benchmark price decline.
    https://epoch.ai/data-insights/llm-inference-price-trends
  4. Axis Intelligence Research: "AI Inference Cost Statistics 2026: The Market That Split in Two"
    Source for the LLMflation Index's mid-2025 peak, its July 2026 reading of 333, and the frontier-versus-budget price divergence.
    https://axis-intelligence.com/ai-inference-cost-statistics/
  5. VoxBooster: "AI Inference Cost Statistics (2026): 50+ Data Points on the Price Collapse"
    Consolidated cross-check of the a16z, Stanford HAI, and Epoch AI figures cited above.
    https://voxbooster.com/blog/ai-inference-cost-statistics-2026/

Published: August 30, 2026. Sources verified at time of publication. All external links open in a new tab.

A small silver Mac mini sitting on a wooden desk at night, its single status light glowing, with a faint abstract network pattern reflected on the surface beside it.

Your Mac Mini Can Run Real AI Models With Zero Cloud Involved — Here's What It Actually Takes.


A stack of legal case law books glowing softly at the edges as if digitized, with a faint neural network pattern spreading across their spines, set against a dark law-office background.

A Million Professionals Already Use This Company's AI. This Week It Started Writing Legal Briefs on Its Own.


A stylized lightning bolt made of glowing code fragments and terminal windows, striking downward against a Google-colored gradient background, with a faint empty silhouette labeled 'Pro' fading in the distance behind it.

Google Just Released Its Third "Flash" AI Model in Six Weeks. Its Actual Flagship Is Still Missing.