Intelligence Falls to $2 a Million Tokens While Micron Posts an 87% Margin
Frontier model prices are converging at $2 per million input tokens, yet the memory that runs them is getting scarcer. Investors may want to ask what becomes scarcer as AI usage grows.

The economics of AI are inverting. API prices for frontier models are converging at $2 per million input tokens, yet the memory and context storage that run them are becoming a bottleneck. The question for investors is less how smart AI has become than what grows scarcer as AI usage rises.
Frontier AI converges on $2
On Sept. 22, OpenAI launched GPT-6 Sol at $2 per million input tokens and $10 for output. On Sept. 28, Anthropic priced Claude Sonnet 5.5 the same way. On Sept. 30, Google set the introductory price of Gemini 4 Argon at $2 and $10 as well.
Price is not the only thing that moved. According to Google, Argon ranked first on the long-horizon software engineering benchmark DeepSWE v1.1 at 77.9% and on Zapier AutomationBench at 51.3%, and it raised the maximum output limit to 1 million tokens. Anthropic says Sonnet 5.5 keeps the same per-token price as the prior generation but needs fewer tokens for the same task, cutting cost per task by up to 30%.
Epoch AI, which tracks the cost of reaching a given level of performance, finds that costs have fallen about 47% per quarter since 2023, or roughly 13x a year. The direction looks less like a temporary promotion than a structural trend.
Meanwhile, Micron posted an 87% gross margin
Software prices are falling, but physical infrastructure is moving the other way. Micron's fiscal fourth-quarter 2026 revenue was $54.229 billion, about 4.8 times the $11.315 billion a year earlier. Non-GAAP EPS was $33.42, and non-GAAP gross margin was 87.0%.
- $54.23B: FY4Q26 revenue, up about 379% from a year earlier
- 87.0%: FY4Q26 non-GAAP gross margin
- $61.5B: midpoint of FY1Q27 revenue guidance
The contract figures stand out more. According to Reuters, customer financial commitments tied to long-term supply agreements rose from $22 billion in June to $32 billion, mostly as cash deposits. Micron's remaining performance obligations grew to about $150 billion. The company expects memory and storage supply to be tighter in fiscal 2027 and 2028 than in 2026, and says it has already secured contracts for most of its 2027 HBM output.
Reading this as "the memory cycle is gone" would be risky. Long-term contracts improve revenue visibility and strengthen the supplier's bargaining power. They do not rule out customer capex adjustments, technology shifts or a downturn once new supply arrives.
Jevons paradox: cheaper AI may need more infrastructure
The mechanism is simple. When intelligence gets cheaper, the same budget buys more calls. A chatbot answers once. An agent plans, searches, runs code and fixes failures in a loop, so token use per task rises sharply.
If token prices halve and usage more than doubles, total infrastructure demand does not fall. That is why the Jevons paradox, in which efficiency gains raise total consumption, could show up in AI. It is an economic hypothesis, not an automatic law. What decides the outcome is how far usage growth outruns the pace of price declines.
The next bottleneck is not only HBM: KV cache and SSDs
As long contexts and multi-turn agents spread, GPUs must do more than compute. They must keep prior context available. That is the job of the KV cache. Its size varies widely with model architecture, precision and the number of KV heads, so it is wrong to say that 1 million tokens always needs 320GB. Still, one research example calculated that a 1-million-token context for Llama 3 70B needs about 320GB of KV cache, far beyond a single 80GB H100.
That pushes designs toward tiering context across CPU DRAM and flash storage rather than holding it all in HBM. NVIDIA presents its BlueField-4-based CMX as a shared context layer that extends GPU memory, and claims up to 5x the throughput and power efficiency of conventional storage.
Micron's earnings materials are notable here. The company says its PCIe Gen 5 7600 SSD and Gen 6 9650 SSD are already shipping to major customers for KV-cache applications. It is a real commercial signal that the investment map for "AI memory" could widen from HBM to server DRAM and enterprise SSDs.
Who benefits across the value chain
| Layer | Positives | Key risk |
|---|---|---|
| Memory and storage | HBM shortage, expanding long-term contracts, demand for server DRAM and SSDs for KV cache | Price spikes may prompt customer optimization and capacity additions |
| Cloud and big tech | Cheaper models drive more usage and cloud traffic | Higher memory, power and network costs squeeze inference margins |
| Vertical AI | Falling model costs for a key input; room for higher-value, outcome-based pricing | Foundation models expanding features and bundling |
| General-purpose model providers | Scale economies and higher usage | Intensifying price competition; token margins compress without differentiation |
Insight Times Editorial Desk





