Investing

If GPUs Are the Muscle, the Network Is the Nervous System

AI data centers do not get faster just by buying more GPUs. When thousands of accelerators cannot exchange data on time, the most expensive asset in the building sits idle. That is the whole investment case for Arista.

Photo Intel Free Press · CC BY 2.0 · Wikimedia Commons

In an AI cluster the GPU is the most expensive thing. A waiting GPU is the most wasteful.

In an AI data center, the most expensive asset is usually the accelerator.

One GPU is expensive. Gather thousands of them and the capital outlay climbs fast.

But that expensive GPU is not always computing.

It waits for data produced by another GPU. It waits for memory. It waits for a network transfer to finish. It waits for a communication collision to clear.

Stretch those waits out and GPU utilization falls.

If you bought 10 billion dollars of GPUs and the real utilization rate is low because of the network, the network bottleneck produces the same result as having bought fewer GPUs.

So as AI clusters grow, the network stops being a peripheral component. It becomes the equipment that sets the actual return on the compute asset.

Why AI is more network-sensitive than ordinary cloud

Traditional cloud workloads were built around relatively independent servers.

AI training is different. One enormous model is split across thousands of GPUs.

What one GPU computes has to reach another GPU before the next operation can run. Gradients are summed, parameters are synchronized, tensors are moved.

Inference needs GPU-to-GPU communication too, once a model is spread across multiple accelerators.

In other words, a single AI job makes thousands of chips behave like one giant computer.

That is why in an AI cluster bandwidth is not the only thing that counts. Latency, congestion, packet loss and load balancing all feed directly into job completion time.

If the network is slow, the system is slow no matter how fast the chip is.

You cannot see Arista until you separate scale-up from scale-out

Scale-up is the domain where accelerators sitting very close together are bound into one giant system. Inside a server, inside a rack, or within very short reach of a rack. It demands low latency and extremely high bandwidth.

NVIDIA holds a strong position here through NVLink and NVSwitch.

Scale-out is wider. It connects racks and clusters, thousands to tens of thousands to hundreds of thousands of accelerators.

This is where the Ethernet ecosystem is strong. It is also where Arista has spent years building an advantage in large-scale Ethernet fabric.

The line between scale-up and scale-out used to be fairly clear. As AI rack-scale architectures get bigger, the two are starting to overlap.

That is why Arista is pushing its 1.6T platform beyond scale-out and into scale-up.

Why Ethernet is getting stronger in AI again

For a long time the story in AI networking was InfiniBand versus Ethernet.

InfiniBand was strong in HPC and in the first wave of large-scale AI training, on low latency and raw performance.

Ethernet's advantages are an open standard, a huge supply ecosystem, a range of switch silicon, an operating model cloud operators already know, and the ability to pick among multiple vendors.

Once a cluster reaches hundreds of thousands of XPUs, operability and supply chain matter as much as performance.

Rather than depending on one vendor's proprietary structure for the entire network, combining accelerators, switches and NICs on an Ethernet base can look attractive to a hyperscaler.

Arista's long-term investment case rests on the claim that Ethernet can deliver enough performance for AI workloads while keeping openness and operational flexibility.

1.6T is not simply twice 800G

1.6T means a single port handles 1.6 terabits of data per second. Twice 800G.

But in an AI data center the meaning is not just double the speed.

The same number of ports connects more GPUs. The same rack space supports a larger fabric. Fewer switch tiers are needed. Fewer network hops means lower latency and lower power draw.

Arista's 7060XE7 line delivers 102.4Tbps of bandwidth per system. Configurations include 64 ports of 1.6T or 128 connections at 800G.

At that level a switch is no longer a connection device. It is part of the rack-scale AI architecture.

The new competitive unit is not Gbps. It is useful compute per GPU.

Networking vendors have traditionally sold port count and speed.

100G, 400G, 800G, 1.6T.

In the AI era those numbers are not enough on their own.

The end result an investor should look at is how much less the GPU waits.

Assume a network improvement lifts GPU utilization from 60 percent to 70 percent. With the same GPU count, usable compute rises by roughly 17 percent.

The real effect varies with workload and topology, but the economic principle holds.

The network can be a small share of server cost. Yet if it improves utilization across the entire GPU asset base by even a few percentage points, the ROI effect is far larger.

Congestion is the traffic jam of the AI data center

When thousands of GPUs transmit at once, traffic does not flow evenly.

At a given moment data piles onto a given path, switch buffers fill, packets are delayed, retransmissions happen.

Hence the importance of congestion management and load balancing in AI networks.

Arista is building AI fabric features into the 7060XE7 and EOS, including dynamic load balancing and cluster load balancing.

The point is not simply to build a fast road. It is to spread traffic efficiently across many paths so that no single stretch becomes the bottleneck.

In an AI system, the slowest segment sets the pace of the whole job.

Arista's real moat is EOS, not the hardware

Seeing Arista as a box vendor is the same category of error as seeing NVIDIA as a GPU chip vendor.

One of Arista's core assets is EOS, the Extensible Operating System.

Networking in a large cloud data center does not end when the switch is installed.

You have to manage the state of thousands or tens of thousands of devices, monitor traffic, find faults, push updates, automate, and hold a consistent policy across the entire cluster.

AI data centers are more complex in both scale and traffic pattern.

So networking software matters more, not less.

Once Arista's EOS is embedded deep in a hyperscaler environment, swapping out the switch hardware stops being a simple decision.

Arista's moat comes from operating software and customer workflow as much as from fast port speeds.

What the Q2 2026 numbers show: not just growth, but margin

Arista posted revenue of 3.036 billion dollars in Q2 2026.

That is up 37.7 percent from 2.205 billion dollars a year earlier, and the first quarter above 3 billion dollars.

GAAP operating margin was 45.4 percent. Non-GAAP operating margin was 49.9 percent.

Non-GAAP EPS rose roughly 40 percent year over year.

These numbers matter because Arista is holding very high margins while growing fast.

That means large cloud customers and data centers are paying up for software, operability and high-performance Ethernet.

But a high margin also means high expectations may already be priced in.

A long-term investor should watch not only the growth rate but how well that margin survives competition and customer negotiation.

Arista's biggest strength is also its biggest risk: the cloud titans

Arista has deep relationships with the world's largest cloud companies.

In the AI era that is its greatest strength.

Hyperscalers move from 400G to 800G first, then to 1.6T. They build the biggest AI clusters.

Arista gets to validate new products with these customers and deploy at scale.

Customer concentration is also a risk.

Large cloud buyers have enormous purchasing power, design their own networks, and can choose white box hardware and their own software.

A schedule change on a single project can swing a supplier's results.

So the biggest risk to Arista's AI growth is not that Ethernet disappears. It is that hyperscalers take more of the network's value in-house.

Is NVIDIA a competitor or a market maker?

Reading the Arista and NVIDIA relationship as straight competition is also inaccurate.

NVIDIA has NVLink, InfiniBand and Spectrum-X. It is a strong competitor in scale-up and across AI networking generally.

At the same time, the more NVIDIA GPUs sell, the bigger the total scale-out network market gets.

NVIDIA is both a threat to Arista and the thing creating its market.

What matters over the long run is which segments of an AI cluster belong to proprietary fabric and which belong to open Ethernet.

Arista's push to extend 1.6T Etherlink into scale-up territory is an attempt to move that boundary inward, bit by bit.

Power and cooling become part of network design

As AI network speeds rise, so does the power problem.

High-speed SerDes, optical transceivers and switch ASICs all consume power and generate heat.

Gather thousands of optical modules and the network itself becomes a serious power draw.

So in the shift to 1.6T, watts per bit matters alongside throughput.

Arista offers both air-cooled and liquid-cooled configurations of the 7060XE7.

It is also pointing to Linear Pluggable Optics, or LPO, as a way to cut related optical power consumption sharply, by the company's own account.

The scope of AI system optimization is widening from the GPU out to the network.

The bigger the cluster, the more the systems companies win

In older data centers you could think of server, network and storage as fairly independent parts.

In AI those boundaries blur.

GPU topology, network topology, memory, power, cooling and software scheduling all act on each other.

Arista itself describes the 7060XE7 not as a switch but as part of a rack-scale system.

The competitive unit in the AI industry is moving up from component to system.

Just as NVIDIA has moved from GPU to rack-scale AI factory, Arista is trying to move from switch vendor to AI fabric system vendor.

Who pulls off that transition may decide long-term margins.

The most dangerous illusion in owning Arista is thinking GPU shipments are all you need to watch

Arista is a beneficiary of AI infrastructure spending.

But NVIDIA GPU shipments and Arista revenue do not move one for one.

What matters is cluster topology.

Where the GPUs sit. How large a single cluster gets. How fast Ethernet port speed climbs. How many tiers the network is built in. How much of their own network stack hyperscalers use. When the 800G to 1.6T transition really begins.

These are the variables that set Arista's revenue and margin.

So the fact that AI GPU capex keeps rising does not complete the Arista thesis by itself.

You have to look separately at how much of that network value Arista actually captures.

The most expensive asset is the GPU, but the most wasteful thing in an AI data center is a GPU that is waiting.

Sources

Insight Times Editorial Desk