Infra, p. 2
Infra
42 stories · page 2 of 5Cloudflare rearranged one record in memory and got 100 terabytes back, and the cache got faster
The platform behind 1.1.1.1 holds over 250 billion cache entries at any moment. Five successive changes to how they are laid out cut each entry by more than half. The memory freed is roughly the entire RAM of 130 of their servers.
Read →NVIDIA moved the memory controller off the chip and put it inside the memory itself
On 26 August NVIDIA added NVHBM to NVLink Fusion - memory whose base die now carries their own controller. The claim is up to 30 per cent more bandwidth, 15 per cent less power and up to a quarter of the compute die freed up. Amazon's Annapurna Labs goes first.
Read →Meta unveils MTIA 300, its first chip built to train models
MTIA 300 is Meta's first accelerator built to train models, not just serve answers from finished ones. It targets recommendation and ranking systems, with the network built straight into the silicon. By Meta's own numbers, total communication time drops 3.9x versus an equivalent GPU cluster.
Read →Meta wrote its own AI network protocol and opened it through the Open Compute Project
MetaRoCE is a transport protocol written from scratch for the conversations between thousands of AI accelerators over plain Ethernet. By the company's own numbers, it keeps about 86 percent of throughput even at one percent packet loss. The spec, the reference code and the compliance test are going public.
Read →NVIDIA reports 30x more agentic throughput per megawatt, but the number still awaits independent review
NVIDIA published its own measurements for Vera Rubin NVL72: 30 times more agentic throughput per megawatt and 35 times lower cost per token versus the previous generation. The numbers are theirs, on a workload they ran themselves, and by the company's own words still await review from SemiAnalysis.
Read →Groq 3 LPX enters production, built together with NVIDIA's platform
A rack-level inference accelerator, co-designed with Vera Rubin NVL72. It combines GPUs with dedicated fast-response modules and handles token generation in agentic systems. Already deploying at Nebius, CoreWeave and SpaceXAI.
Read →AWS Glue 6.0 cuts the price by 30 percent and adds full Iceberg v3 support
AWS shipped version 6.0 of Glue, its data-processing service, at 30 percent lower cost than previous versions. Underneath is a new engine, and on top, full support for the Iceberg table format v3.
Read →GitHub explained the 17 August outage: capacity did not keep up with its own growth
On 20 August GitHub published a post-mortem of the 17 August failure, which held the platform down for seven hours and 47 minutes and took out github.com, authentication, Actions, APIs and Copilot. Per the company the cause was not a change in code or configuration but capacity that did not arrive in time.
Read →Cerebras launched CS-4 and promised up to thirty times faster than GPU systems
The fourth generation of their system stands on three processors the size of a whole silicon wafer. The company claims up to 30 times faster inference. The number is theirs; there is no independent check yet.
Read →