In MLPerf Inference v6.1 NVIDIA submitted the first Vera Rubin NVL72 results against GB300 NVL72: up to 3.7 times higher throughput on Qwen3-VL and up to 2.5 times on DeepSeek-R1. It is a shared test with rules, but the system sits in the category for upcoming products.
- The comparison is against NVIDIA's own previous generation, on models it chose to submit.
- On DeepSeek-R1 in the offline scenario, four GB300 NVL72 racks reach 99 per cent scaling efficiency against one.
- The 30 times figure from SemiAnalysis's AgentX test is NVIDIA's own claim and sits outside MLPerf.
Three weeks ago NVIDIA gave its own number for Vera Rubin, measured on its own ground. We wrote then that it was waiting for an outside check.
Now comes the first thing that looks like one. Not quite, but closer.
What preview means
MLPerf is a shared exam with shared rules: the same models, the same scenarios, checking by the organiser. That is why its numbers weigh more than a slide. The preview category, however, is for systems that are not on sale yet, so the exam is real and the product is not yet on the shelf.
And second: against whom it is measured. 3.7 times is against their own previous generation, on a model and with settings NVIDIA itself chose to submit. The competitor is not in the sentence at all.
The 99 per cent scaling figure for the current generation is the most boring and the most useful of all. In this test four racks give almost four times what one does, meaning the money for the next rack is not thrown away. For whoever pays the power bill, that is the line to read first.