September 27, 2026 · by Geoff
Token economics: the ceiling stays high, the audience shrinks, and the floor goes to zero
The price of a fixed capability falls 13x a year while the price of the newest frontier model has risen 5x in twelve months. That is not a contradiction. It is the shape of the market, and it tells you who ends up owning the frontier.
machine version: /machine/blog/token-economics.md
This is the argument behind the Token Economics report, a page of charts and a dataset that updates every week. The essay changes rarely; the numbers change often. Every figure here is in the report with a source.
Two prices for the same token
There are two price series for language models and they are moving in opposite directions.
The first is the price of a fixed level of capability. What GPT-4 could do in March 2023 cost $37.50 per million tokens. By February 2025 the same level of capability cost $0.18. By September 2026 the cheapest model that clears the bar costs about a dime. Epoch AI, which measures this across five benchmarks, puts the decline at 47% per quarter, roughly 13x a year, faster than compute, faster than DNA sequencing, faster than batteries ever were. A 75% score on GPQA Diamond cost thirty cents a question with o3 and four ten-thousandths of a cent eighteen months later. Reaching o3's score on ARC-AGI-1 cost $4,560 a task in December 2024 and thirty cents by August 2026.
The second series is the price of the newest frontier model, and it is U-shaped. OpenAI's top tier fell from $60 per million output tokens to a floor of $10 with GPT-5 in August 2025. Then it rose at every release: $14, $15, $30, $30, and $50 for GPT-6 Astra in September 2026, with a fast mode at $100. Anthropic cut Opus threefold, then put Fable and Mythos above it at $50. Google relaunched the Flash name at 25 to 30 times its 2024 price. DeepSeek, the lab that started the 2025 price war, raised its Pro tier by as much as 1,100%. Every provider that sells a frontier re-established a tier above its previous top inside twelve months.
Neither series is lying. The cost to serve a token is low and falling: bottom-up estimates put marginal cost near $3 per million output tokens on rented H100s against $15 to $50 list, SemiAnalysis puts Anthropic's inference margin above 70% in the Blackwell era, and Google's TPU v7 serves a 400-billion-parameter model for eighteen cents per million all-in. The rising ceiling is a pricing decision made under scarcity. Inference demand is growing about 10x a year against compute supply growing about 3.4x, and when capacity is scarce and your best customers are price-insensitive, you price the newest model at what the top 1% will pay and let last year's capability fall to cost. That is exactly what the charts show.
The audience for the very top shrinks as it improves
Here is the claim I care most about, because it is where the money goes next. As models get better, the audience for the very top tier gets smaller, not larger, and the ceiling stays high anyway.
The adoption data already says so. One month after launch, Claude Fable 5 was 6% of Anthropic tokens and 11% of Anthropic spend on Ramp's card data, while the mainstream flagship took the volume. GPT-6 Astra is gated behind a $200-a-month plan and OpenAI paused new sign-ups for it under demand. The top 1% of customers generate 80% of revenue at both OpenAI and Anthropic. OpenRouter's price elasticity is nearly zero: a 10% price cut buys 0.5% more usage. These are the fingerprints of a specialised product with a narrow, deep-pocketed buyer, not a mass one.
Why do those buyers pay? Because of what they are doing with it. Anthropic's own data shows that Claude Code, which runs long agentic sessions, uses Opus 54% of the time, while chat uses it 10% of the time. Practitioners have a rule of thumb: an autonomous loop upgrades the model one tier, a human in the loop downgrades it one tier. The arithmetic is unforgiving. A step that succeeds 90% of the time is 57% reliable over eight steps. Ninety percent end-to-end on a ten-step workflow needs 99% per step. When nobody is watching, the extra points at the top are the whole product.
That is the threshold. Today, businesses default to the top tier for anything unattended because no cheaper model can be trusted to run for long without someone checking. METR's time-horizon work, the best measurement we have, puts the frontier's 50%-success horizon at around a working day in early 2026, doubling roughly every four months. But 50% is not a colleague. METR's own note says reliability-critical tasks need 98% or better to be worth automating, the 80%-success horizons are about five times shorter, and nothing above sixteen hours is measurable on the current suite. OpenAI's September 2026 "research intern" milestone needed human intervention in more than half of its successful four-to-eight-hour tasks.
So the threshold is not crossed, by anyone, at any price. The moment it is, the picture flips. When a mid-tier model can run for months as a genuine digital double for an employee, the default stops being "buy the best" and becomes "buy the cheapest thing that clears the bar". Epoch's data says the lag from a capability debuting at the frontier to costing 10x less is four to eleven months. Once "reliable enough to leave alone" is a capability like any other, it will follow the same curve. The frontier will keep its price by moving the target: longer horizons, harder evaluations, reasoning tokens billed by the thousand, $200 and $300 subscriptions that sell outcomes rather than tokens. Its audience will be the people whose problems genuinely need it, and there are fewer of those every quarter, not more.
Two honest complications. Enterprise buyers also cite talent scarcity and data security, not only reliability, when they choose closed models, and surveys attribute most agent failures to integration and infrastructure rather than model reasoning. And substitution is already happening at the request level, through routers rather than through trust: Cursor's router reports frontier-quality results at 60% lower cost by sending only the hard turns to the expensive model. The frontier loses most of the requests and keeps most of the spend, for now.
The floor goes to zero, but not the way the West expected
The bottom of the market is already at hardware cost. OpenAI's own open-weight gpt-oss-120b is hosted by two dozen providers from $0.03 per million input tokens; a dozen of them charge an identical $0.15/$0.60, which is what a commodity looks like. Every one of OpenRouter's top ten models by weekly volume lists input at $0.84 or less. A gaming GPU you already own serves batched tokens at 28 to 40 cents per million at full utilisation and beats a $1.70 API above about 15% utilisation; the calculator on the report page lets you plug in your own numbers. Gemma 4 runs on a Raspberry Pi. Gemini Nano runs on a phone at 19 tokens a second. Apple's on-device model activates one to four billion parameters and does summarisation, classification and tool calls without leaving the device.
Specialisation pushes the floor lower still. Fine-tuned models under eight billion parameters beat GPT-4 on most of 31 narrow tasks in the LoRA Land study, each on a single consumer GPU. NVIDIA's researchers argue small models are the right default for most agent sub-tasks at a tenth to a thirtieth of the cost. Products like TypeSafe AI's Jev, a non-generative decision model that answers classification, routing and gating questions with typed outputs, take a different route to the same place: carve the well-understood operations out of the general model entirely, perfect them, and stop paying frontier prices for them. The vendor's cost claims are its own and one early tester found it more expensive than Gemini, so treat it as a direction rather than a result. The direction is right.
What I did not expect is who set the floor. Enterprise open-source share fell from 19% to 11% between 2024 and 2025, Meta released its first closed model in April 2026, and the true open frontier, DeepSeek, Qwen, Kimi, GLM, is Chinese and about four months behind the closed frontier. The bottom fell out, and the labs that pulled it out are in Hangzhou and Beijing. That matters for the last part.
The frontier gets nationalised, quietly
The obvious prediction is that the state takes over the frontier labs. I think that is wrong in form and right in substance, and the data shows the form it actually takes.
The capital wall comes first. Four hyperscalers spent about $141 billion on capex in 2023, $360 billion in 2025, and have guided $725 to $800 billion for 2026 including Oracle. Alphabet posted its first negative free-cash-flow quarter and raised $85 billion of equity. OpenAI raised $122 billion at $852 billion; Anthropic raised $65 billion at $965 billion; both explicitly for compute. A frontier training run alone was projected to cross $1 billion by 2027, and that is the small number. Nobody enters this market with venture capital. Every 2022-vintage challenger has been absorbed through a licence-and-hire deal, Inflection, Adept, Character, Windsurf, or has pivoted, and the new entrants exist as satellites of Nvidia, SpaceX or the incumbents. Antitrust reviewed all of it and cleared all of it.
Ownership comes second, and it is not the state. Sovereign funds sit on every cap table: SoftBank at 13% of OpenAI with MGX and Amazon and Nvidia behind it, GIC and Temasek and MGX in Anthropic, the Qatari, Saudi and Omani funds in xAI, the European Commission's fund and Bpifrance in Mistral. The infrastructure layer is a single $100 billion vehicle that ties BlackRock, Microsoft, MGX, Nvidia and xAI together. On the public side the index managers hold 8 to 20% of every large-cap AI company, though founders keep control through dual-class shares. I looked for documented ties between the asset managers and intelligence agencies and found none; what is documented is more direct than that.
Because the third channel is procurement and coercion, and it is where the state actually shows its hand. In 2025 the Pentagon put four labs on $200 million prototype contracts and the GSA put ChatGPT and Claude into federal agencies for a dollar each. When Anthropic insisted on two limits, no fully autonomous weapons and no domestic mass surveillance, the government ordered every agency to stop using it and designated it a supply-chain risk, cutting a lab valued near a trillion dollars out of the defence-contractor market. A district court called that unlawful retaliation; the appeals court upheld the designation in September 2026. The equity template exists too: a 9.9% stake in Intel, a 15% cut of Nvidia's China revenue, a golden share in US Steel, and a reported proposal from OpenAI itself to hand 5% to a public wealth fund as political insurance. The state does not need to own a lab to control it, and the labs know it.
Regulation is the fourth channel, and the smallest, except in Europe. The EU's rules for general-purpose models presume systemic risk above a compute threshold that only a handful of companies cross, and put the obligations on those companies. Mistral lobbied to soften the rules, signed the code of practice when it came, joined calls to pause the Act, took government money in its last two rounds, holds French and Luxembourg military contracts, and by most accounts trails the frontier by a year or more. That is what a regulatory ladder pulled up behind a champion looks like. In the United States, regulation of model development is falling, not rising.
Put the four together and the frontier of 2030 is not a nationalised industry. It is a handful of companies that cannot exist without hyperscaler capital, sovereign money and government procurement, that lose their largest customer if they set conditions on how they are used, and that sit behind a capital wall no newcomer can climb. Governance will arrive as a condition of the money, not as a law. The door will close without anyone announcing it closed.
What to watch
Three numbers on the report page will tell you when the story turns. The 80%-success horizon: when it reaches a work-week, the threshold argument becomes a schedule, and the mid tier clears the same bar within a year. The lag from frontier debut to a 10x cheaper equivalent: four to eleven months today, and if it shortens, the frontier's premium window closes with it. And the first government equity stake in a frontier lab, which would turn "entangled" into "owned".
The tokens themselves will keep getting cheaper. The question was never whether. It is who gets to sell the ones that still cost money.