On 30 September NVIDIA said CoreWeave had made Vera Rubin NVL72 systems available and that Cognition, the company behind Devin, is the first customer running real work on them. In early tests Cognition saw up to 4.8 times higher total token throughput for SWE-2 compared with GB200 NVL72. The figures are NVIDIA's and the customer's, not independent.
- Cognition measured on a sample of FrontierCode tasks solved by agents.
- The Vera CPU comes to CoreWeave with 11,264 cores per rack - more than 11,000 isolated agent environments.
- CoreWeave also launched Forge, an environment for training and improving models and agents.
Four point eight times. The number is a customer's, even though NVIDIA publishes it, and that makes a difference.
Cognition, who make Devin, have not been selling Vera Rubin to anyone. They took tasks from FrontierCode, set agents to solve them and compared against the previous GB200 generation. One customer's test on one workload, but at least it is their workload.
Part of the weight is in the CPU. An agent needs a sandbox, and reinforcement learning needs thousands of sandboxes at once. A rack with more than eleven thousand isolated environments answers exactly that need.
The line about the V100s is nice too: the old cards still work nearly a decade on. It is advertising, of course, buy today and it will run for years. But it is also a rare reminder that old iron does not die on the day the new one comes out.
The next number has to come from someone the seller is not quoting.