The Number That Matters More Than How Smart AI Gets

This week's AI models got stronger and cheaper at the same time. The chart long-term investors should watch isn't the weekly benchmark leader, it's the falling cost of finishing a task successfully.

This week's news isn't a model swap. It's a price curve.

On September 22, Anthropic released Claude Opus 5.5. The company says it runs on typical workloads at roughly 40% lower cost than Opus 5. API pricing dropped to $4 per million input tokens and $20 per million output tokens, each 20% below the previous Opus 5.

The same day, OpenAI launched GPT-6 Sol and Luna. GPT-6 Sol's price fell from $4 to $2 per million input tokens and from $20 to $10 per million output tokens. Luna dropped from $0.20 to $0.10 on input and from $1.20 to $0.50 on output. OpenAI frames this as a 50% cut versus GPT-5.6's promotional pricing.

xAI's Grok 4.7 arrived on September 21. Pricing held steady at $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, while performance improved on long-context coding and knowledge work. The pattern across all three is clear: newer models aren't getting more expensive. They're doing more for the same price, or the same work for less.

CLAUDE OPUS 5.5 -40% — Anthropic's stated reduction in typical operating cost versus Opus 5

GPT-6 SOL -50% — API input and output price cut versus GPT-5.6 Sol

GROK 4.7 $2 / $6 — Per million input / output tokens. Same price as Grok 4.6, with improved performance

A number that outlasts the benchmark: an Intelligence Price Index

The AI industry produces a new benchmark leader almost every week. But the more useful question for long-term investors isn't "who's on top," it's "how much does a given level of intelligence cost to buy?"

Call it, informally, an Intelligence Price Index. It isn't an official industry metric, but the trend is already being tracked. According to the Stanford AI Index, the inference cost to reach GPT-3.5-level performance on the MMLU benchmark fell from $20 per million tokens in November 2022 to $0.07 in October 2024, a drop of more than 280-fold in about 18 months.

Token price alone isn't enough, though. Reasoning models vary in how many tokens they burn to reach an answer, and agents retry after failures. So the metric the industry actually needs to watch is Cost per Successful Task, the total cost of completing one piece of work successfully. Artificial Analysis has started comparing frontier models this way too, ranking cost per task under its Intelligence Index rather than by raw token price.

A better way to measure AI economics

AI cost per successful task = price per call × tokens actually used + retry cost + tool-use cost

When the price hits $1, the boundary of automation moves

Falling prices don't just cut a company's AI bill. They widen the range of tasks that make economic sense to automate.

If AI handles a task for $100, a human worker may still be cheaper. At $10, AI starts competing on some repetitive tasks. At $1, a much larger set of processes becomes a candidate for automation. Below $0.10, some tasks may become uneconomical to hand to a person at all.

  • AI cost per task: $100 — human labor still has the edge in most cases
  • AI cost per task: $10 — competition begins on repetitive tasks
  • AI cost per task: $1 — large-scale automation becomes economically attractive
  • AI cost per task: $0.10 — always-on agents become a realistic option

These figures are illustrative, not an industry-wide average cost. The point is the threshold. Companies don't just ask whether AI is "smarter" than a human. They ask whether it's accurate enough, and cheaper, faster, and scalable around the clock.

Is cheaper AI bad news for Nvidia?

Intuitively, yes. If the same answer takes less compute, it looks like less GPU demand is needed. But if demand grows faster than prices fall, the result flips. It resembles what economists call the Jevons Paradox.

Imagine an agent that used to run 10 times a day because of cost now running 1,000 times a day instead, reviewing every code commit, drafting a personalized reply for every customer, generating every ad individually, and continuously screening every financial transaction.

The paradox of AI demand

Cost per intelligence ↓ Total intelligence consumed ↑↑

So a cut in model prices should not be read directly as a cut in compute demand. What matters is which moves faster: the price decline or the growth in usage. If usage grows faster, the price per model call falls even as the total volume of inference a data center must handle keeps rising.

Seen this way, the long-term case for Nvidia shifts slightly too. Training is concentrated computation that builds a giant model once. Inference is the repeated computation that runs every time users and agents actually use the finished model. As AI moves deeper into everyday workflows, what investors need to watch isn't only the size of training clusters, but the volume of everyday inference calls and how fast that volume is growing.

What to watch after the token price

MetricWhat it measuresWhy it matters to investors
Price per TokenThe sticker price on an API callEasy to compare, but misses differences in tokens used per model and retry costs
Cost per Successful TaskThe total cost of actually completing one taskThe closest proxy to real agent economics
AI Cost / Human Labor CostAI cost for a task versus the cost of a human doing itTasks where this ratio falls below 1 see a stronger automation incentive
Inference VolumeActual call volume, tokens, and total computeShows whether falling prices are shrinking or growing GPU demand
The metric to track now isn't how smart a model is, it's how fast the cost of one successful task keeps falling.

Insight Times Editorial Desk