AI's Next KPI Isn't Token Price. It's Cost Per Successful Task

Apple is cutting the cost of where inference happens, Anthropic is cutting the cost of running its models, and Snorkel AI shows money flowing into the data and evaluation layer that reduces failed attempts. All three point to one shift: AI competition is moving from bigger models to cheaper successful outcomes.

The new unit of AI isn't the token. It's the completed task.

The AI cost a company actually pays is never just the API price sheet. When a model fails to finish a task in one pass, it burns more tokens, calls tools again, needs a human to fix the output, and picks up extra security, data-integration and operating costs along the way.

Cost per successful task = total cost of model, infrastructure, review and data ÷ number of tasks completed correctly

At the same price per million tokens, a model that finishes a job in fewer steps, with a lower failure rate and less human intervention, is actually far cheaper.

That means the numbers worth tracking going forward are Cost per Successful Task, Time to Completion, Human Intervention Rate, and the revenue or labor savings generated per dollar spent on AI. These metrics are likely to become the common language for comparing models, cloud providers, chipmakers and enterprise software companies against each other.

The chain runs in four steps:

  1. Data and reinforcement learning — get the model to reason better from the start, cutting the failure rate.
  2. The model — finish the same task using fewer tokens and fewer inference steps.
  3. Execution infrastructure — run inference wherever it's cheapest: cloud, local, or edge.
  4. Workflow — turn the cost savings into real labor automation and revenue.

Snorkel AI: after GPUs, the next bottleneck is good problems and good grading

Snorkel AI announced a $350 million Series E on September 22, putting its valuation at $3.5 billion. The more striking number is revenue. The company says it has grown more than 18-fold since launching its data-services business in 2025, pushing its annualized revenue run rate past $375 million.

That number isn't a comeback story for data labeling. As frontier models push into coding, math, law, medicine and science, where correct answers are hard to generate and even harder to verify, what's needed isn't a simple answer key. It's reward functions, rubrics and reinforcement-learning environments designed to judge whether a complex task was actually done right.

If a model is more accurate from the start, inference tokens, tool calls, rollbacks and human rework all fall together. That is where Snorkel sits on the AI cost curve. Still, investors should not treat the run rate as locked-in annual revenue. Customer concentration, how much of that service revenue depends on human labor, and how much converts into higher-margin software revenue matter more.

Anthropic: a 40% drop in task cost matters more than a 20% price cut

ItemPriceNote
Input tokens$4 / 1Mdown 20% from Opus 5's $5
Output tokens$20 / 1Mdown 20% from Opus 5's $25
Cache reads$0.20 / 1Mdown 60% from $0.50

Anthropic says Claude Opus 5.5 runs at roughly 40% lower cost than Opus 5 on typical workloads with default settings. The price sheet came down 20%, but the cost of getting a job done fell 40%. By the company's account, that's because the new model finishes the same work using fewer tokens.

In the agent era, that gap matters. For long-running coding sessions or workflow automation, cost isn't set mainly by the price of processing one prompt. It's set by how long the agent works, how many times it fails, and how much context it has to re-read. That's why the 60% cut in cache-read pricing matters more for agentic workloads than for ordinary chat.

Still, the 40% figure comes from Anthropic's own testing. Enterprise customers would need to run their own A/B tests against their codebase, documents, toolchain and security requirements, and measure success rates and human review time, not just take the vendor's number at face value.

Apple: it's changing the price of where inference happens, not the price of the model

Apple and LM Studio, at WWDC 2026, ran a 1-trillion-parameter Kimi K2.6 model across four Mac Studios linked by Thunderbolt 5 and RDMA. In August, Apple followed up with the M5 Ultra Mac Studio, officially supporting up to 512GB of unified memory per unit and clustering across multiple Mac Studios.

That's not a claim that this replaces Nvidia data centers. Large-scale training and services used simultaneously by thousands or tens of thousands of users still favor data-center GPUs. What Apple appears to be targeting realistically is enterprise inference with high repeat usage, data that's hard to move off-site, and comparatively low concurrent access.

Upfront hardware cost + power and management costs vs. recurring API costs + cloud GPU rental + data-transfer and security costs

Local AI doesn't make tokens free. It shifts variable cost into fixed cost centered on hardware depreciation. The economics improve as utilization rises, and cloud stays the better option when usage is low or swings sharply.

Where the investment map is shifting

Investment axisHow it makes moneyRepresentative companiesKPIs to watch
Compute and memoryRising AI usage expands total compute demandNvidia, Broadcom, AMD, TSMC, Micron, and othersInference volume, HBM capacity, AI capex, network demand
CloudSells deployment, management, security and data services togetherMicrosoft, Amazon, Alphabet, OracleAI revenue, cloud growth rate, capex payback, inference margin
ModelsCompletes the same task with fewer tokens, less time, lower failure rateAnthropic, OpenAI, Google, xAICost per successful task, cache utilization, customer retention
Data and RLCuts failures and retries through post-training and evaluation environmentsSnorkel AI and specialized data vendorsGrowth rate, gross margin, customer concentration, contract length
Enterprise AIMonetizes real workflow automation and labor cost savingsMicrosoft, ServiceNow, Palantir, Salesforce, and othersAI ARR, automation rate, human intervention rate, revenue per seat
Local and edgeConverts repeat inference into fixed hardware costApple, PC OEMs, related chipmakersHigh-capacity memory mix, utilization rate, enterprise references, TCO
AI cost is set not by token price but by the odds a task finishes correctly on the first try.

Insight Times Editorial Desk