SANTA CLARA, 30 AUG 2026 — Nvidia's Vera Rubin platform has ramped into full production. The headline claims are 30 times higher throughput per megawatt and 35 times lower token costs — measured against Nvidia's own GB300 NVL72, on metrics Nvidia chose.
What is shipping
Vera Rubin pairs the Rubin GPU with the Vera CPU at rack and pod scale. Nvidia describes a POD-scale platform of five purpose-built racks operating as a single system for agentic workloads, with the NVL72 rack integrating liquid cooling, power smoothing, a cable-free MGX architecture and hot-swappable NVLink switch trays.
Spectrum-X Ethernet Photonics is also in production, combining co-packaged optics with Spectrum-X switching, which Nvidia positions as the path to million-GPU deployments.
We covered the Vera CPU's architecture when it was presented at Hot Chips, including the decision to put every core on one of its six dies. This is the point at which those design choices become deployed hardware.
How to read a generational efficiency claim
None of these numbers is a like-for-like comparison with a competitor. They compare a new Nvidia rack to an older Nvidia rack, on workloads Nvidia selected, using efficiency metrics Nvidia defined.
Nothing about that is dishonest or unusual. Every silicon vendor reports generational gains this way, and comparing against your own previous product has a decent claim to being the fairest available test, since you control both sides of it. It does mean the figures answer the question "how much better is this than what we sold you last year" and not "how does this compare to the alternatives".
The gap between a 30x per-megawatt gain and a 10x per-unit-energy gain against Blackwell is also instructive. Different baselines and different metrics produce different multiples for the same hardware, which is a reason to read the comparison clause rather than the number.
We applied the same discipline when OpenAI's Jalapeño beat Blackwell on work per watt, which was the expected result — a purpose-built inference chip should beat a general-purpose GPU on inference efficiency, and saying so is not a criticism of either.
What a POD-scale product tells you about the buyer
The unit of sale is worth pausing on. Nvidia is not selling a card or a server here; it is selling five racks that operate as a single machine.
That tells you who the customer is. An enterprise does not buy a five-rack minimum with liquid cooling, integrated power smoothing and a specific networking fabric to install in an existing hall. It requires a facility designed around it, which means the purchasing decision sits with hyperscalers, neoclouds and national AI programmes rather than with corporate IT.
It also raises the switching cost considerably. A buyer standardising on a cable-free MGX architecture with NVLink switch trays and Spectrum-X photonics has committed the building, not only the compute. Competing silicon has to displace a facility design, not a part number.
For regional operators this narrows the field of who can host frontier hardware at all. A colocation provider whose halls were built for air cooling and 15kW racks is not in this market, regardless of its floor space.
Why per-megawatt is now the metric that matters
The shift from raw performance to performance per megawatt is the substantive story, and it is being driven by grid capacity rather than by electricity bills.
For data centre operators across this region the binding limit is what a utility will connect, not what the electricity costs once connected. Malaysia's utility reports a pipeline of 8.3GW against about 1.05GW of load actually drawn, and Malaysia's own forecast has data centres taking 31 per cent of national electricity by 2035.
In that setting, hardware that does more work per megawatt is not a cost optimisation. It is more compute inside a fixed connection agreement, which is the only variable an operator with an approved load can still move.
What the efficiency gain does to demand forecasts
There is a consequence here that coverage of efficiency and coverage of grid capacity rarely join up, and it runs against the intuition.
Every regional electricity forecast for data centres is built on assumptions about how much power a given amount of AI compute requires. A large generational efficiency gain should reduce the power needed per unit of work, which would make those forecasts too high.
Historically it has not worked that way. Cheaper compute per watt has reliably produced more compute rather than less power, because the constraint moves to what the workload can profitably consume, and AI workloads have shown no sign of saturating. The efficiency gain is more likely to raise total demand by making previously uneconomic work worth doing.
A regulator reading a 30x claim should therefore not revise a load forecast downwards. The better response is to ask operators what they intend to do with the headroom, since the historical answer has been to use it.
What to watch
Independent measurement, and specifically MLPerf results with published power figures, which is where vendor claims in this category have historically been tested against a common methodology.
Also worth watching is whether any regional operator publicly attributes an increase in contracted capacity to per-megawatt gains. That would be the first evidence that efficiency is translating into connection decisions rather than into marketing, and it would show up in a utility's filings before it showed up in a press release.