The mode runs on NVIDIA Blackwell and is available in the OpenAI API and to eligible ChatGPT Work and Codex users. OpenAI has no announcement of its own in its newsroom; the number is in NVIDIA's blog, with quotes from two OpenAI people, and in OpenAI's API documentation. OpenAI's documentation adds the condition that, for us, comes first: US or global processing only, no EU endpoints.
- 8x is against Astra Standard, on token generation. The number is NVIDIA's and also stands in OpenAI's documentation; no third party has measured it.
- Ultrafast is the fastest and more expensive service tier of the API: OpenAI says to use it when speed justifies the cost. For GPT-5.6 Sol it is in preview.
- No EU processing. If you hold data that must not leave Europe, this mode is not for you.
The agent writes code, calls a tool, reads the result, decides. And again. Hundreds of times for one task. Every second of waiting in a loop like that gets multiplied by hundreds, which is why news about speed weighs more than news about one more capability.
The number 8 is a vendor number and we write it as one. It comes from NVIDIA, which sells the chips, and stands in the documentation of OpenAI, which sells the tokens. Two sellers, one number. I am not saying it is wrong. I am saying nobody outside has measured it and that "up to 8x" means "sometimes less". Measure it yourself on your own load before you quote it to a client.
The fine print is elsewhere: Ultrafast supports US data residency and global processing only, with no European endpoints. If you hold client data that must not leave the region, this mode is not for you, however fast it is. That is written by OpenAI in its documentation, not by NVIDIA on its blog.
What comes next: the price. OpenAI calls it higher and points to a table we are not quoting until we have compared it with Standard on the same task. Whether Ultrafast is for agents in production or for demos is decided by the price of a finished task, not the price of a token.