Why Nvidia Is Happy to Help Make AI Models Nearly Free
Falling model prices can squeeze model makers. For Nvidia, the number that matters is the total volume of tokens generated worldwide and the compute that produces them.

Cheap models do not make compute free
According to Axios, Reflection is preparing a powerful open-weight model. Its early performance may trail the top US closed models. The target is to compete with the leading Chinese open-weight family. The key point is that a company can take the model weights and run them on its own data and infrastructure.
Nvidia's interests differ somewhat from those of model companies. OpenAI and Anthropic sell intelligence through APIs and subscriptions. Nvidia sells the computing infrastructure that trains and runs that intelligence. Which model ranks first matters less than how many models and agents run computation, and how often.
Nvidia wants less a single best model than a market in which the total amount of AI to be computed keeps growing.
Reflection already shows the structure
Ironically, Reflection is buying vast amounts of compute to build "AI that can be used cheaply." According to Reuters, the company has a contract with SpaceX to use computing capacity based on Nvidia GB300 systems. It pays $150 million a month starting in July 2026. Annualized simply, that is about $1.8 billion.
A low price for an open model does not erase the cost of building it or running it at scale. If more companies and countries want to own their models, GPU demand now hidden inside API providers could spread across more buyers. Those include enterprise AI factories, sovereign AI programs and neoclouds.
| Metric | Figure |
|---|---|
| Reflection's monthly payment under the SpaceX compute deal | $150M |
| Nvidia FY2027 Q2 Data Center revenue | $89.0B |
| Year-over-year growth in Data Center revenue, same quarter | +117% |
Free models can generate demand for Nvidia
Nvidia's Nemotron strategy is a good clue. Nvidia releases not only model weights but also some training data and recipes, and it allows commercial deployment. Running these models safely and at scale in a real enterprise setting, however, can involve paid layers: NIM microservices, AI Enterprise, GPU systems and networking.
This differs a little from the razor-and-blades model. Nvidia gives models away at nearly no cost to create more reasons to compute. When the computing happens, it tries to capture economic value in hardware and systems software.
Model prices fall → AI and agent use rises → total compute grows
There is a resemblance to the Jevons Paradox in economics. If efficiency per unit of compute improves and token prices fall, but usage grows faster still, total resource consumption can rise. When Nvidia unveiled Rubin in fiscal 2026, it set a goal of cutting inference token costs to as little as one-tenth of Blackwell's. In effect, Nvidia designs its chips to make tokens cheaper while expecting far more tokens to be produced.
Not a formula that cheaper AI always favors Nvidia
First, open-weight models could become efficient so quickly that the GPU time needed for the same work drops sharply. Second, if companies keep concentrating on clouds such as AWS, Azure and Google Cloud rather than building their own, the number of demand sources may grow but Nvidia's customer concentration problem remains. Nvidia disclosed that in FY2026 two direct customers accounted for 22% and 14% of total revenue.
Third, if the growth in AI usage from lower prices fails to outpace efficiency gains, the Jevons effect weakens. Fourth, if alternative accelerators such as Google's TPU, Amazon's Trainium and AMD's chips gain share, the equation "more AI use means more Nvidia use" could break.
So the number for investors to watch is not the model price itself. It is the product of three things: how fast the price per token falls, how fast total token generation grows, and the share of that computation that Nvidia's platform captures.
Insight Times Editorial Desk





