The Number That Matters More Than How Smart AI Gets
This week's AI models got stronger and cheaper at the same time. The chart long-term investors should watch isn't the weekly benchmark leader, it's the falling cost of finishing a task successfully.

This week's news isn't a model swap. It's a price curve.
On September 22, Anthropic released Claude Opus 5.5. The company says it runs on typical workloads at roughly 40% lower cost than Opus 5. API pricing dropped to $4 per million input tokens and $20 per million output tokens, each 20% below the previous Opus 5.
The same day, OpenAI launched GPT-6 Sol and Luna. GPT-6 Sol's price fell from $4 to $2 per million input tokens and from $20 to $10 per million output tokens. Luna dropped from $0.20 to $0.10 on input and from $1.20 to $0.50 on output. OpenAI frames this as a 50% cut versus GPT-5.6's promotional pricing.
xAI's Grok 4.7 arrived on September 21. Pricing held steady at $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, while performance improved on long-context coding and knowledge work. The pattern across all three is clear: newer models aren't getting more expensive. They're doing more for the same price, or the same work for less.
CLAUDE OPUS 5.5 -40% — Anthropic's stated reduction in typical operating cost versus Opus 5
GPT-6 SOL -50% — API input and output price cut versus GPT-5.6 Sol
GROK 4.7 $2 / $6 — Per million input / output tokens. Same price as Grok 4.6, with improved performance
A number that outlasts the benchmark: an Intelligence Price Index
The AI industry produces a new benchmark leader almost every week. But the more useful question for long-term investors isn't "who's on top," it's "how much does a given level of intelligence cost to buy?"
Call it, informally, an Intelligence Price Index. It isn't an official industry metric, but the trend is already being tracked. According to the Stanford AI Index, the inference cost to reach GPT-3.5-level performance on the MMLU benchmark fell from $20 per million tokens in November 2022 to $0.07 in October 2024, a drop of more than 280-fold in about 18 months.
Token price alone isn't enough, though. Reasoning models vary in how many tokens they burn to reach an answer, and agents retry after failures. So the metric the industry actually needs to watch is Cost per Successful Task, the total cost of completing one piece of work successfully. Artificial Analysis has started comparing frontier models this way too, ranking cost per task under its Intelligence Index rather than by raw token price.
A better way to measure AI economics
AI cost per successful task = price per call × tokens actually used + retry cost + tool-use cost
When the price hits $1, the boundary of automation moves
Falling prices don't just cut a company's AI bill. They widen the range of tasks that make economic sense to automate.
If AI handles a task for $100, a human worker may still be cheaper. At $10, AI starts competing on some repetitive tasks. At $1, a much larger set of processes becomes a candidate for automation. Below $0.10, some tasks may become uneconomical to hand to a person at all.
- AI cost per task: $100 — human labor still has the edge in most cases
- AI cost per task: $10 — competition begins on repetitive tasks
- AI cost per task: $1 — large-scale automation becomes economically attractive
- AI cost per task: $0.10 — always-on agents become a realistic option
These figures are illustrative, not an industry-wide average cost. The point is the threshold. Companies don't just ask whether AI is "smarter" than a human. They ask whether it's accurate enough, and cheaper, faster, and scalable around the clock.
Is cheaper AI bad news for Nvidia?
Intuitively, yes. If the same answer takes less compute, it looks like less GPU demand is needed. But if demand grows faster than prices fall, the result flips. It resembles what economists call the Jevons Paradox.
Imagine an agent that used to run 10 times a day because of cost now running 1,000 times a day instead, reviewing every code commit, drafting a personalized reply for every customer, generating every ad individually, and continuously screening every financial transaction.
The paradox of AI demand
Cost per intelligence ↓ Total intelligence consumed ↑↑
So a cut in model prices should not be read directly as a cut in compute demand. What matters is which moves faster: the price decline or the growth in usage. If usage grows faster, the price per model call falls even as the total volume of inference a data center must handle keeps rising.
Seen this way, the long-term case for Nvidia shifts slightly too. Training is concentrated computation that builds a giant model once. Inference is the repeated computation that runs every time users and agents actually use the finished model. As AI moves deeper into everyday workflows, what investors need to watch isn't only the size of training clusters, but the volume of everyday inference calls and how fast that volume is growing.
What to watch after the token price
| Metric | What it measures | Why it matters to investors |
|---|---|---|
| Price per Token | The sticker price on an API call | Easy to compare, but misses differences in tokens used per model and retry costs |
| Cost per Successful Task | The total cost of actually completing one task | The closest proxy to real agent economics |
| AI Cost / Human Labor Cost | AI cost for a task versus the cost of a human doing it | Tasks where this ratio falls below 1 see a stronger automation incentive |
| Inference Volume | Actual call volume, tokens, and total compute | Shows whether falling prices are shrinking or growing GPU demand |
Insight Times Editorial Desk





