NVIDIA published its own measurements for Vera Rubin NVL72: 30 times more agentic throughput per megawatt and 35 times lower cost per token versus the previous generation. The numbers are theirs, on a workload they ran themselves, and by the company's own words still await review from SemiAnalysis.
- By NVIDIA's own numbers: 30x more agentic throughput per megawatt versus GB300 NVL72.
- Cost per token drops 35x, again by their own measurements.
- Measured by NVIDIA itself on the SemiAnalysis AgentX workload, and by the company's own words the SemiAnalysis review has not happened yet.
I read the number twice before moving on. Thirty times more agentic work for every wasted megawatt. Then I saw the line under it, small type: measured by NVIDIA, awaiting review from SemiAnalysis.
Whose number is this
The review has not happened. Until it does, thirty times is a measurement from the company that sells the hardware, on a workload it chose to run through it itself. That does not make it false. It means nobody outside has repeated it yet.
There is one more place where the number is incomplete. It does not count what the Vera processor does when the agent calls tools mid-task, and that is exactly what an agent does half the time. You cannot rerun their test. You can wait for the review, and only then put the number into your own math.