OpenAI published the first measurements of its own inference chip. They claim higher throughput and shorter waiting at the same time, which until now was treated as a trade. The numbers are theirs and were measured by them.
- Jalapeño is OpenAI's first custom inference chip, and these are its first published results.
- The claim is throughput and latency together, not one at the cost of the other.
- The chip was designed with AI, and designed so that AI could program it.
Announcing a chip is the easy part. The numbers come later and are usually more modest than the stage.
These came, and they are not modest.
Why this is not the usual boasting
Inference carries an old trade. Fill the machine with more simultaneous requests and you get more work out, but every single person waits longer. Optimise for the fast answer and the machine sits half empty.
Anyone who has served a model knows that compromise and picks a side. The claim here is that you do not have to pick.
The sentence above is theirs and it is more interesting than the benchmarks. The circle closes: models help build the hardware they will later run on, and the hardware is deliberately shaped so a model can handle programming it.
How I read it
As a claim, not a result. A vendor's numbers about a vendor's product are the start of a conversation, not the end. They become fact when someone outside runs the same models on the same hardware and gets the same thing.
But the direction is clear and it is not news only for OpenAI. The big players no longer buy compute, they build it, because the electricity bill decides who survives faster than the bill for cards does.