TL;DR
- On September 1, 2026, Silicon Data's LLM Token Expenditure Index fell to $0.97 per million tokens — the first time it has dropped below $1. - This is less than half the summer peak, and down roughly 97% from GPT-4's launch price of $60 per million tokens in 2023. - The fall is driven by Chinese open-source models, OpenAI price cuts, and hardware efficiency gains — not a drop in demand. - For businesses using AI, this changes what's now economically viable to build. Agent workflows, large-scale document processing, and always-on AI features that were cost-prohibitive in 2024 can now be run profitably. - For AI providers, it creates a margin problem. Token deflation compresses revenue while compute infrastructure costs stay fixed. - The businesses that benefit most are those charging for outcomes, not tokens.
A million AI tokens now costs less than a dollar.
On September 1, 2026, Silicon Data's LLM Token Expenditure Index — a usage-weighted benchmark that tracks what the market actually pays across real model traffic, listed on Bloomberg under SDLLMTK — fell to $0.97. It was the first time the index had crossed below $1 since its launch, and less than half the reading from its summer peak.
To understand why that matters, it helps to know where this started. In March 2023, a million output tokens from GPT-4 cost roughly $60. By early September 2026, you can run a frontier-capable model for under $1 per million input tokens. That is a 97% reduction in roughly three years. And by all available evidence, the curve is not flattening.
What drove the latest drop
The September 1 reading wasn't a single event. It reflects several forces converging at once.
The most direct factor is Chinese open-source models. Moonshot's Kimi K3, DeepSeek V4, and Alibaba's Qwen 3.6 have reached capability levels that compete with GPT-4-class models on standard benchmarks — coding, summarization, general reasoning — at prices well below what Western providers charge. On major multi-provider routing platforms, the four most-used models at certain points this year were all Chinese. That volume shift pulls the usage-weighted average down even before anyone cuts list prices.
OpenAI accelerated things further on July 30, cutting GPT-5.6 Luna prices by 80% and Terra prices by 20%. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens — pricing that would have been unimaginable for a production-grade model two years ago.
Underlying all of this are hardware and software efficiency gains. The move from Hopper to Blackwell GPU architecture, combined with FP8 and FP4 quantization, smarter batching, and speculative decoding, means providers can serve more tokens from the same hardware. When serving costs fall, competitive pressure forces the savings to pass through to customers.
Steve Hou, Silicon Data's head of research, put it plainly: the combined supply of frontier and lower-cost models may already be sufficient to handle most tasks. Supply has caught up with demand. In a commoditizing market, prices compress.
The price curve matters more than the price
The 97-cent reading is significant. The trajectory behind it is more important.
AI inference costs have fallen at roughly 10x per year for equivalent capability since 2023. That rate is faster than Moore's Law and faster than any comparable cost curve in enterprise software. Guido Appenzeller at a16z tracked what it costs to run a model of constant quality over time. A GPT-3 quality model cost $60 per million tokens at launch. By late 2024 it cost $0.06. That's a 1,000x reduction in roughly three years.
If the 10x-per-year curve continues, a million tokens will cost roughly $0.10 by mid-2027. Tasks that barely pencil out at current prices become trivially cheap. Workflows that require running agents over thousands of documents, analyzing every customer interaction in real-time, or generating personalized content at scale move from pilot projects to standard operations.
The businesses already planning for that curve are in a different position from the ones reacting to each price announcement.
What this changes about what you can build
Three categories of AI use cases shift meaningfully when tokens drop below $1 per million.
High-volume document processing. Running AI over every contract in a legal discovery, every support ticket in a backlog, or every product review a retailer receives was cost-prohibitive at $10–$60 per million tokens. At $0.97, the math changes. A company processing 10 million tokens per day pays under $30. The same workload cost $600 eighteen months ago.
Always-on agent workflows. AI agents use significantly more tokens than simple chat interactions — typically 5 to 30 times more per completed task because of iterative reasoning, tool use, and context re-transmission. That's why agent deployments stayed expensive even as chat prices fell. At sub-$1 token prices, running agents continuously on business processes — monitoring, summarization, triage, drafting — moves from expensive experiment to viable product feature.
Smaller companies catching up. Until recently, the economics of sophisticated AI features favored large enterprises that could negotiate volume contracts and absorb high per-token costs. At current prices, a ten-person startup running 50 million tokens per month pays under $50. The capability gap between well-resourced and under-resourced teams is narrowing faster than anyone expected.
The model routing decision
One implication most teams haven't fully internalized: using a single AI provider for everything is no longer the right approach.
The price spread between models has widened, not narrowed. GPT-5.6 Luna at $0.20/M input and Claude Fable 5 at $10/M input are both in active production use. A team that routes all traffic to Fable 5 out of habit is spending 50x more than necessary on tasks that Luna handles adequately.
Optimal routing in 2026 means segmenting by task complexity. Simple extraction, classification, and summarization tasks go to budget-tier models: DeepSeek V4 ($0.30/M), Gemini Flash-Lite ($0.25/M), or GPT-5.6 Luna ($0.20/M). Longer context tasks and multi-step reasoning go to mid-tier. Only complex, high-stakes work — legal analysis, agentic reasoning chains, nuanced synthesis — justifies frontier-tier pricing.
Teams implementing task-based routing typically cut their AI inference spend by 40–70% while maintaining the same output quality across their products. The savings compound quickly at any meaningful usage volume.
The provider problem
Not everyone benefits from falling token prices. For the companies selling model access, the dynamics are uncomfortable.
Charles-Henry Monchau, chief investment officer at Syz Group, described the structural issue directly: "Token deflation compresses the revenue line while compute infrastructure commitments stay fixed." Servers are leased or owned. Power contracts run for years. Staff are employed. When per-unit revenue falls faster than those costs, margins get squeezed — sometimes severely.
This is a particularly acute problem for OpenAI and Anthropic, both of which filed confidentially for IPOs this summer. Investors in public markets care about pricing power and the long-term unit economics of AI. Sub-$1 tokens raise legitimate questions about whether foundation model companies can build durable, high-margin businesses or whether they will function more like utilities — essential, large-scale, but not particularly profitable.
Monchau's view on how providers respond is worth reading closely: "The strategic response is visible: moat must shift away from raw capability — where the open-weight gap is now measured in months — to distribution, memory and context." In other words, model quality alone is no longer a sustainable competitive advantage. What matters is who has the largest installed base of users, the richest context about those users, and the deepest integration into workflows people run daily.
That's why OpenAI is investing in ChatGPT's memory features. It's why Anthropic is building deep enterprise integrations. Raw intelligence is becoming a commodity. The platform built around it is not.
The business model that survives deflation
Token deflation is not inherently good or bad for AI businesses. It depends entirely on how you are priced.
If your product charges customers per token or per API call, falling prices compress your revenue directly. Every provider price cut is a revenue cut. Every efficiency gain you pass through reduces your margin. You are exposed to the commodity curve in both directions.
If your product charges for outcomes — per contract reviewed, per support ticket resolved, per lead qualified, per document summarized — falling token prices are pure upside. Your cost of goods sold drops. Your price to customers holds. Margins expand.
Intercom charges $0.99 per resolved support ticket. Salesforce Agentforce charges $2 per conversation. Neither of those prices is pegged to inference costs. When tokens get cheaper, those companies profit more per unit, not less.
The exercise worth running right now: if inference were free tomorrow, would anyone still pay for your product? If yes, you have built something with value beyond the AI layer. If not, your product is the inference, and a free-inference world means a free version of your product exists.
What to do with the price drop
For product teams and founders, falling token prices create three practical opportunities.
First, revisit features you shelved as too expensive. The use cases that didn't pencil out at $5–$10 per million tokens may now make sense to build. Run the numbers again with current pricing.
Second, implement routing if you haven't already. Most teams running a single premium model on all tasks are overspending by a factor of three to ten. Basic task-complexity segmentation recovers that spend without touching product quality.
Third, think harder about pricing structure. If you're selling on a per-token or per-call basis, start modeling what happens when prices fall another 10x. Build your business around the work you do and the outcomes you deliver, not the compute you consume.
The token price curve isn't slowing down. The businesses that outperform over the next two years will be the ones that treat it as a structural condition to plan around, not a quarterly variable to monitor and react to.
At 97 cents per million tokens, AI intelligence is no longer a scarce resource. The constraint is no longer cost. It's knowing what to do with it.



