we_are_coded.by CODE · The world, decoded
БГ
NVIDIA

Vera Rubin entered MLPerf at up to 3.7 times the previous generation, for now as a preview result

NVIDIA Blog · event date: 16 September 2026Infra

In MLPerf Inference v6.1 NVIDIA submitted the first Vera Rubin NVL72 results against GB300 NVL72: up to 3.7 times higher throughput on Qwen3-VL and up to 2.5 times on DeepSeek-R1. It is a shared test with rules, but the system sits in the category for upcoming products.

In short
  • The comparison is against NVIDIA's own previous generation, on models it chose to submit.
  • On DeepSeek-R1 in the offline scenario, four GB300 NVL72 racks reach 99 per cent scaling efficiency against one.
  • The 30 times figure from SemiAnalysis's AgentX test is NVIDIA's own claim and sits outside MLPerf.
Checked on1 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

Three weeks ago NVIDIA gave its own number for Vera Rubin, measured on its own ground. We wrote then that it was waiting for an outside check.

Now comes the first thing that looks like one. Not quite, but closer.

The facts: on 16 September 2026 MLCommons released the MLPerf Inference v6.1 results. Per NVIDIA, in its first preview submission Vera Rubin NVL72 delivers up to 3.7 times higher throughput than GB300 NVL72 on Qwen3-VL, across the offline, server and interactive scenarios, with vLLM and NVIDIA Dynamo, and up to 2.5 times on DeepSeek-R1 with TensorRT-LLM. The results are in the Closed Division, entries 6.1-0106 and 6.1-0074. A DeepSeek-R1 submission with 288 GPUs in four GB300 NVL72 racks reached 99 per cent scaling efficiency in the offline scenario against a single rack, and software optimisations raised GB300 NVL72 performance on Qwen3-VL by up to 1.6 times over v6.0. Nebius also submitted preview results for Vera Rubin NVL72. In the same post NVIDIA repeats its own figure from the SemiAnalysis AgentX test - 30 times better performance than GB300 NVL72 in preview testing.

What preview means

MLPerf is a shared exam with shared rules: the same models, the same scenarios, checking by the organiser. That is why its numbers weigh more than a slide. The preview category, however, is for systems that are not on sale yet, so the exam is real and the product is not yet on the shelf.

And second: against whom it is measured. 3.7 times is against their own previous generation, on a model and with settings NVIDIA itself chose to submit. The competitor is not in the sentence at all.

Fast against yourself is good news for NVIDIA. Against everyone else, we do not know yet.

The 99 per cent scaling figure for the current generation is the most boring and the most useful of all. In this test four racks give almost four times what one does, meaning the money for the next rack is not thrown away. For whoever pays the power bill, that is the line to read first.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→NVIDIA Blog - NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut, 16.09.2026→MLCommons - MLPerf Submission Rules (Preview Systems)
Original: https://wearecoded.com/en/articles/vera-rubin-mlperf-3-7-pati-gb300.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news