AI's Biggest Power Draw May Not Be Training. It Could Be Everyday Use
Anthropic's 2.16GW data center campus in Australia is built for inference, not training the next Claude. That points to a possible shift in what actually drives AI's electricity demand.
Two nuclear reactors' worth of power, spent on work, not study
2.16GW — planned maximum capacity of the Western Downs Digital Park 18.9TWh — theoretical annual maximum if that 2.16GW ran nonstop, 24/7 Inference — not training a new model, but processing real user requests to Claude
The number 2.16 gigawatts doesn't mean much on its own. It's a bit more than the combined rated output of two large 1-gigawatt nuclear reactors. If that capacity ran continuously, 24 hours a day, 365 days a year, the theoretical maximum annual consumption would be about 18.9 terawatt-hours. Actual usage will depend on utilization rates, phased buildouts, and power efficiency.
Australia's public broadcaster ABC reported that Western Downs Digital Park could use roughly 47 gigawatt-hours a day once fully built out, a massive load compared with Queensland's current average daily consumption of 170 gigawatt-hours. It's worth noting the project is still at the proposal stage, and construction could take years.
But the more important word here isn't the power figure. It's inference. Training is the process where an AI learns how to solve problems. Inference is what happens after that: a trained model receiving a real customer's request and generating an answer. Put simply, if training is studying, inference is showing up to work every day.
Streaming that's heavier than Netflix
Imagine a movie that cost $200 million to make. That production process is like training. Streaming the finished film to hundreds of millions of people worldwide requires CDNs and servers. That's closer to inference.
But AI is far more computationally expensive than Netflix. Netflix delivers a video file that already exists. Generative AI computes new tokens every time a question comes in. Reasoning models run longer internal computations before answering. Agents plan, search, run code, check results, and retry when they fail.
What looks like a single request to a user can trigger multiple model calls inside a data center. In its 2026 report, the International Energy Agency noted that newer AI usage patterns like reasoning and agentic tasks can consume hundreds to thousands of times more energy per request than simple text generation. At the same time, model and hardware efficiency keep improving fast, so the ultimate power demand hinges on which moves faster: efficiency gains per request, or the sheer surge in usage.
The formula for AI compute is changing
In generative AI's early days, investors focused on training. Bigger models, more GPUs, longer training runs — that was AI infrastructure investing.
But once a service goes mainstream, the economics change. Training is very expensive but concentrated at specific points in time. Inference happens every time a customer uses the product. As user counts grow, questions get longer, and agents break a single request into multiple steps, total compute demand accumulates.
Early stage: Training-centric → Mid stage: Training + Inference → Mature stage: Massive Inference + Agentic Workloads
So the center of gravity in AI compute may be shifting from "training-centric" to "training plus large-scale inference," and eventually toward "always-on agentic workloads." Anthropic's Australia deal is an early signal that inference is turning from a secondary task into a business that requires locking up gigawatt-scale power infrastructure in advance.
That said, this alone doesn't prove inference already consumes more power than training. Anthropic hasn't disclosed actual GPU counts per facility, utilization rates, power consumption, or the overall split between training and inference compute. What this news shows isn't that the reversal has already happened. It's that infrastructure investment premised on that reversal is starting.
For Nvidia investors: a bigger market, and a different kind of competition
This shift is both good news and a challenge for Nvidia. If inference demand explodes, the total addressable market for GPUs, networking, HBM, and power and cooling infrastructure could expand. But the metrics customers care about are also changing.
In training, top-line performance and the ability to scale massive clusters mattered most. In inference, profitability comes down to tokens per dollar, tokens per watt, latency, and utilization. When Nvidia describes its 2026 Blackwell Ultra chips, the numbers it leads with aren't "how fast is one GPU" but "how many tokens per megawatt" and "cost per token."
That's exactly where competition is intensifying. Custom silicon like Google's TPUs and Amazon's Trainium chips, along with AMD accelerators, don't need to beat Nvidia at everything. If they can serve a specific model cheaply and efficiently enough, they can take a slice of the inference market. So the growth of inference widens Nvidia's total market, but it also creates a market where being the fastest chip alone isn't enough.
Insight Times Editorial Desk





