Token Economics

What a million tokens costs, and why the ceiling and the floor are moving in opposite directions

The price of a fixed level of capability falls about 13x a year. The price of the newest frontier model bottomed in August 2025 and has risen since. Open-weight models run on a gaming GPU for cents. This page tracks all of it from primary sources, updates weekly, and ships every series as JSON.

Edition 2 · data as of 2026-09-27 · updated weekly · /api/token-economics

Read the essay: the ceiling stays high, the audience shrinks, the floor goes to zero →

This edition

2026-09-27

Verification pass run the same day as edition 1's launch; no material change to any series. Checked provider pricing pages, Artificial Analysis, OpenRouter, and METR for anything published since edition 1 and found only confirmations of figures already recorded, so every series is left as-is this edition.

  • OpenAI GPT-6 Astra ($10/$50, launched 2026-09-03), GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50), both launched 2026-09-22, already match the recorded tierLadder (https://developers.openai.com/api/docs/pricing).
  • Anthropic Claude Fable 5.1 ($10/$50, launched 2026-09-01) and Claude Opus 5.5 ($4/$20 flagship tier, launched 2026-09-22) already match the recorded frontier and tierLadder rows (https://platform.claude.com/docs/en/about-claude/pricing).
  • Google Gemini 3.8 Flash intro pricing ($0.75/$3.75 input/output through 2026-12-31, reverting to $1.50/$7.50) already matches the recorded tierLadder note (https://ai.google.dev/gemini-api/docs/pricing).
  • 2026 hyperscaler capex guidance reconfirmed in the same range already on file: Amazon ~$200B, Alphabet $195-205B (one report cites 'as much as $205B'), Meta $130-145B (https://www.globaldatacenterhub.com/p/microsoft-q3-fy2026-the-190b-capex, https://valueaddvc.com/blog/meta-145b-ai-capex-2026-why-zuckerberg-raised-guidance-twice).
  • METR's Time Horizon 1.1 methodology (the source already cited for metrHorizons/metrNotes) remains the latest published update; no new dated model horizon point has been released since edition 1 (https://metr.org/blog/2026-1-29-time-horizon-1-1/).
  • No new open-weight leaderboard score, OpenRouter volume figure, or enterprise usage marker was found beyond what edition 1 already recorded.

Two speeds: the frontier vs a fixed capability

Output price of the newest OpenAI frontier model against the price of GPT-4-level capability from whoever sells it cheapest. Same tokens, opposite directions.

Sources: epoch.ai, developers.openai.com

Newest frontier price by provider

Each provider's newest top non-Pro model, list price per million tokens, on a log scale. Prices fell for two years, then every provider added a tier above its previous top.

Sources: developers.openai.com, platform.claude.com, ai.google.dev, docs.x.ai, api-docs.deepseek.com, mistral.ai

The price ladder today

Every tier each provider sells right now, per million tokens on a log scale. Read the spread between the top dot and the bottom dot: that is the premium for the newest model.

Sources: developers.openai.com, platform.claude.com, ai.google.dev

How fast a fixed capability gets cheap

Measured declines in the price of reaching a given benchmark score. Epoch AI puts the composite rate at 47% per quarter, about 13x a year.

Sources: a16z.com, hai.stanford.edu, epoch.ai, epoch.ai, r40.io, epoch.ai

Who buys the top tier, and how much of it

The markers behind the demand argument: agents run on the flagship, the ultra-premium tier adopts slowly, and a few hundred customers pay for everything.

Opus share of Claude Code sessions vs chat
54% vs 10%
Jun 2026 · anthropic.com
Fable 5 share of Anthropic tokens / spend on Ramp, one month after launch
6% / 11.4%
Jul 2026 · ramp.com
Top 1% of customers' share of revenue (OpenAI, Anthropic)
80%
2026 · pymnts.com
Open-source share of enterprise LLM usage
19% (2024) → 11% (2025)
2025 · menlovc.com
US labs' share of OpenRouter tokens
~70% (Jun 2025) → ~30% (Jun 2026)
Jun 2026 · pro.stockalarm.io
Cursor Router cost per commit: Balance / Opus 4.8 / Fable 5
$4.63 / $7.34 / $12.69
Jul 2026 · cursor.com
OpenRouter price elasticity
+0.5-0.7% usage per 10% price cut
2025 · openrouter.ai
Enterprise deployments that are true agents
16%
2025 · menlovc.com

Anthropic run-rate revenue, $B

METR 50%-success time horizon

How long a software task a model completes with 50% success, on METR's suite, log scale. The dashed line doubles every 131 days. Hollow points are secondary estimates. Nothing here measures a work-week, and 80%-success horizons are about 5x shorter.

* secondary estimate

Sources: metr.org

Open weights vs closed models: price for score

Artificial Analysis intelligence index against blended price per million tokens (log). Open models sit five to fifteen points behind at a tenth to a fiftieth of the price.

Sources: artificialanalysis.ai

What the floor costs

Hosted open-model prices and self-hosted marginal costs, September 2026. A dozen hosts charge the same $0.15/$0.60 for the same model, which is what a commodity market looks like.

SetupinputoutputNote
gpt-oss-120b, cheapest hosts (AkashML, CoreWeave, FlexAI)$0.03$0.1724 providers listed
gpt-oss-120b, modal price (Together, Groq, Nebius, Azure, Bedrock, OpenRouter)$0.15$0.6
gpt-oss-120b, Cerebras$0.35$0.75~1,700-3,000 tok/s
Llama 3.1 8B (Groq)$0.05$0.08Apr 2026
Llama 3.3 70B (Groq)$0.59$0.79Apr 2026
DeepSeek V4.1 Flash, official (peak / off-peak)$0.3$1.20off-peak $0.15 / $0.60
DeepSeek V4 Pro, official (peak / off-peak)$1.32$3.96off-peak $0.66 / $1.98
GLM-5.3 Flash (OpenRouter)$0.15$0.5#1 by weekly tokens, Sep 2026
Datacenter TCO, 397B MoE on TPU v7 (SemiAnalysis)$0.181total $/M at 100 tok/s per user; B200 $0.222, B300 $0.276
Owned RTX 4090, Qwen 3.8 27B, batched, 100% utilisation$0.28$1.03 at 20% utilisation
Rented RTX 5090 ($0.76/h), Llama 3.1 8B, batch 32$0.06

Self-hosting calculator

Cost per million output tokens on a GPU you own, as a function of how busy you keep it. Hardware is amortised over three years; power is charged per token. Drag the sliders.

Your cost per million output tokens

$0.72/M

Beats the API above 12% utilisation.

RTX 4090 · Qwen 3.8 27B, 16 concurrent requests · 300 W · 3 y

Sources: paraloncloud.com

Hyperscaler capital expenditure

Cash capex by company and year, with 2026 as guidance. The capital wall is the real barrier to entry.

USD billions. Microsoft on a June fiscal year for 2023-2025 and a calendar guide for 2026; Oracle on a May fiscal year shifted one year; 2026 Oracle not captured.

Sources: insight.factset.com

Token volumes

Company-reported throughput on a log scale. Volume grows 5-7x a year while prices at the bottom fall, which is what a commodity looks like.

Google, trillions of tokens per month

OpenRouter, trillions of tokens per week

Sources: blog.google, investing.com, teamday.ai

Changelog

  1. Edition 2 · 2026-09-27

    Verification pass run the same day as edition 1's launch; no material change to any series. Checked provider pricing pages, Artificial Analysis, OpenRouter, and METR for anything published since edition 1 and found only confirmations of figures already recorded, so every series is left as-is this edition.

    • OpenAI GPT-6 Astra ($10/$50, launched 2026-09-03), GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50), both launched 2026-09-22, already match the recorded tierLadder (https://developers.openai.com/api/docs/pricing).
    • Anthropic Claude Fable 5.1 ($10/$50, launched 2026-09-01) and Claude Opus 5.5 ($4/$20 flagship tier, launched 2026-09-22) already match the recorded frontier and tierLadder rows (https://platform.claude.com/docs/en/about-claude/pricing).
    • Google Gemini 3.8 Flash intro pricing ($0.75/$3.75 input/output through 2026-12-31, reverting to $1.50/$7.50) already matches the recorded tierLadder note (https://ai.google.dev/gemini-api/docs/pricing).
    • 2026 hyperscaler capex guidance reconfirmed in the same range already on file: Amazon ~$200B, Alphabet $195-205B (one report cites 'as much as $205B'), Meta $130-145B (https://www.globaldatacenterhub.com/p/microsoft-q3-fy2026-the-190b-capex, https://valueaddvc.com/blog/meta-145b-ai-capex-2026-why-zuckerberg-raised-guidance-twice).
    • METR's Time Horizon 1.1 methodology (the source already cited for metrHorizons/metrNotes) remains the latest published update; no new dated model horizon point has been released since edition 1 (https://metr.org/blog/2026-1-29-time-horizon-1-1/).
    • No new open-weight leaderboard score, OpenRouter volume figure, or enterprise usage marker was found beyond what edition 1 already recorded.
  2. Edition 1 · 2026-09-27

    First edition. Six research strands compiled into the dataset: frontier and mid-tier price history, price-to-capability decline, hyperscaler capex, METR horizons, token volumes, open-weight and self-hosted prices, and enterprise usage markers.

    • GPT-6 Astra ($10/$50) and Opus 5.5 ($4/$20) added as the newest frontier prices; Fable 5.1 unchanged at $10/$50.
    • DeepSeek V4 Pro peak pricing ($1.32/$3.96) recorded after the August increase.
    • OpenRouter weekly volume ~134T tokens (17-22 Sep 2026), top ten all at $0.84 input or less.
    • Capex 2026 guidance: Microsoft ~$190B, Alphabet $195-205B, Amazon ~$200B, Meta $130-145B.

Method: prices are list prices per million tokens from provider pricing pages and launch posts; capex is from earnings releases and filings; horizons are METR time-horizon 1.1 measurements; volumes are company-reported. Figures marked secondary come from press or third-party trackers when the primary page could not be fetched. Every chart has a table view and every figure a source.