Top AI Models Converge at $2 per Million Tokens as Micron Margin Hits 87%
Google, OpenAI and Anthropic now price flagship models alike, while Micron reports an 87% non-GAAP gross margin. The scarce resource is shifting from tokens to memory, storage and power.

Flagship model price lists now point to the same number
Google unveiled Gemini 4 Argon on Sept. 30 at an introductory price of $2 per million input tokens and $10 for output. OpenAI's GPT-6 Sol and Anthropic's Claude Sonnet 5.5 carry the same price. Argon scored 77.9% on DeepSWE v1.1 and 51.3% on AutomationBench, both published by Google. This is hard to read as a discount on a weak model.
The broader trend points the same way. Epoch AI estimates the real cost of reaching a given level of performance has fallen about 47% per quarter since 2023, or roughly 13x a year. Silicon Data's broad LLM token spending index fell to $0.96 per million tokens as of Sept. 24. "Intelligence deflation" shows up in wide market data, not just in marketing copy.
| Metric | Figure | Note |
|---|---|---|
| Frontier API | $2 / $10 | Price per million input / output tokens |
| Epoch AI | About 13x a year | Estimated annual decline in cost for equal performance |
| Silicon Data | $0.96 | Broad-market spending index per million tokens, Sept. 24 |
Yet Micron's gross margin is 87%
Micron's fiscal fourth-quarter 2026 results, released the same day, sent the opposite price signal. Revenue reached $54.229 billion, up 379% from $11.315 billion a year earlier. Non-GAAP EPS was $33.42 and non-GAAP gross margin was 87.0%. Guidance for the next quarter calls for revenue of $61.5 billion, plus or minus $1.5 billion, non-GAAP EPS of $38.15, plus or minus $1, and non-GAAP gross margin of about 86.25%.
The contract numbers matter more. Micron's long-term strategic customer agreements rose to 26, with remaining contract value cited by the market at about $150 billion. Customer cash deposits and related financing commitments grew to $32 billion from $22 billion in June. In the past, volumes were renegotiated on short cycles. Some demand now appears to be locking into long-term contracts.
Concluding that "the memory cycle is gone" would go too far. An 87% gross margin is evidence of severe supply shortage. It also gives customers a stronger incentive to optimize harder and diversify suppliers. The longer high prices last, the more likely new supply and substitute technologies are to follow.
Workload elasticity matters more than Jevons
A halving of model prices does not automatically double infrastructure demand. The key is whether new workloads created by lower prices outrun gains in model efficiency and hardware performance. If they do, total compute and memory use rise.
Model price falls → more agent calls → longer context and repeated runs → more memory, storage and power demand
That is why agent-era economics is hard to explain with unit token prices alone. In work that repeats planning, search, tool calls and error correction, a pricier model that succeeds in fewer iterations can cost less in the end. Anthropic's description of Sonnet 5.5 as "up to 30% cheaper for most tasks" stresses cost per task rather than the API sticker price.
The next bottleneck is not only HBM
As long contexts and long-running agents spread, it gets harder to keep KV cache in GPU HBM alone. Inference infrastructure such as LMCache and Mooncake already supports tiering KV cache from GPU to CPU memory, local SSD and remote storage for reuse. In Mooncake's public benchmark, SSD offloading cut average time to first token by 57% and raised input throughput 2.4x versus GPU-only.
Micron's fiscal fourth-quarter materials also state that its PCIe Gen5 and Gen6 SSDs for KV-cache applications are shipping to major customers. That does not mean SSDs replace HBM. HBM stays the top tier for real-time compute, while cheap, high-capacity SSDs hold the caches and context that must persist.
| Tier | Main resource | Role | What investors should watch |
|---|---|---|---|
| Tier 0 | GPU HBM | Real-time attention, active tokens | HBM capacity, bandwidth, packaging supply |
| Tier 1 | Host DRAM / CXL | Fast cache buffering and swap | Server DRAM bit demand, CXL adoption |
| Tier 2 | Enterprise NVMe SSD | Retaining and reusing large KV cache | QLC eSSD shipments, KV-cache SSD revenue |
| Physical layer | Power, transmission, gas, land | Running the data center | Power timelines and project delays |
A software price war meets hardware scarcity
Project Jupiter shows the last constraint is not only chips. Reuters reported that Oracle issued a force majeure notice over possible delays in securing power for its New Mexico data center. Oracle and STACK said the project is proceeding on its planned schedule. What is confirmed so far is that power and permitting can disrupt the financial timeline of a large AI project.
Opposing forces now operate at once in the AI value chain. Model makers must sell cheaper intelligence. Hyperscalers must handle more requests. Infrastructure suppliers try to keep prices high on limited memory and power. Where that balance breaks, the next investment opportunity and the next risk are likely to emerge together.
Insight Times Editorial Desk





