we_are_coded.by CODE · The world, decoded
БГ
NVIDIA

NVIDIA reports 30x more agentic throughput per megawatt, but the number still awaits independent review

NVIDIA BlogInfra

NVIDIA published its own measurements for Vera Rubin NVL72: 30 times more agentic throughput per megawatt and 35 times lower cost per token versus the previous generation. The numbers are theirs, on a workload they ran themselves, and by the company's own words still await review from SemiAnalysis.

In short
  • By NVIDIA's own numbers: 30x more agentic throughput per megawatt versus GB300 NVL72.
  • Cost per token drops 35x, again by their own measurements.
  • Measured by NVIDIA itself on the SemiAnalysis AgentX workload, and by the company's own words the SemiAnalysis review has not happened yet.
Checked on25 August 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

I read the number twice before moving on. Thirty times more agentic work for every wasted megawatt. Then I saw the line under it, small type: measured by NVIDIA, awaiting review from SemiAnalysis.

The facts: for the Vera Rubin NVL72 platform, versus the previous GB300 NVL72 generation, NVIDIA reports 30x higher throughput per megawatt and 35x lower cost per token on agentic workloads. For comparison, GB300 NVL72 itself had delivered 15x better throughput per megawatt versus the Hopper generation. The measurements were run by NVIDIA on the SemiAnalysis AgentX workload, built from recorded real agentic-coding sessions, and by the company's own words still await review from SemiAnalysis. The results for now do not reflect the work of the Vera processor during tool calls. Source: the NVIDIA blog.

Whose number is this

The review has not happened. Until it does, thirty times is a measurement from the company that sells the hardware, on a workload it chose to run through it itself. That does not make it false. It means nobody outside has repeated it yet.

Thirty times better sounds like a fact. For now it is a manufacturer's claim about its own product.

There is one more place where the number is incomplete. It does not count what the Vera processor does when the agent calls tools mid-task, and that is exactly what an agent does half the time. You cannot rerun their test. You can wait for the review, and only then put the number into your own math.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→NVIDIA Blog: Vera Rubin NVL72 efficiency
Original: https://wearecoded.com/en/articles/nvidia-vera-rubin-nvl72-30x-na-megavat.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news