we_are_coded.by CODE · The world, decoded
БГ
OpenAI

GPT-6 Astra now has an Ultrafast mode: up to 8x faster generation, per NVIDIA

NVIDIAFrontier

The mode runs on NVIDIA Blackwell and is available in the OpenAI API and to eligible ChatGPT Work and Codex users. OpenAI has no announcement of its own in its newsroom; the number is in NVIDIA's blog, with quotes from two OpenAI people, and in OpenAI's API documentation. OpenAI's documentation adds the condition that, for us, comes first: US or global processing only, no EU endpoints.

In short
  • 8x is against Astra Standard, on token generation. The number is NVIDIA's and also stands in OpenAI's documentation; no third party has measured it.
  • Ultrafast is the fastest and more expensive service tier of the API: OpenAI says to use it when speed justifies the cost. For GPT-5.6 Sol it is in preview.
  • No EU processing. If you hold data that must not leave Europe, this mode is not for you.
Checked on2 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

The agent writes code, calls a tool, reads the result, decides. And again. Hundreds of times for one task. Every second of waiting in a loop like that gets multiplied by hundreds, which is why news about speed weighs more than news about one more capability.

The facts: on 1 October 2026 NVIDIA published on its blog that GPT-6 Astra Ultrafast runs on NVIDIA Blackwell GPUs and is available in the OpenAI API and to eligible ChatGPT Work and Codex users. Per NVIDIA, Ultrafast offers up to 8x faster token generation than the Astra Standard mode. The post quotes Philippe Tillet, inference lead at OpenAI, and Uday Ruddarraju, chief technology officer of compute at OpenAI, who says OpenAI used its internal models to optimise inference on NVIDIA GPUs. OpenAI has no announcement of its own in its newsroom. OpenAI's documentation describes Ultrafast as the fastest service tier in the API, up to 8x faster than Standard, broadly available for GPT-6 Astra and in preview for GPT-5.6 Sol, at a higher cost, with a recommendation to use WebSockets for many tool calls in quick succession, and with US data residency and global processing only, with no EU or other non-US regional endpoints.

The number 8 is a vendor number and we write it as one. It comes from NVIDIA, which sells the chips, and stands in the documentation of OpenAI, which sells the tokens. Two sellers, one number. I am not saying it is wrong. I am saying nobody outside has measured it and that "up to 8x" means "sometimes less". Measure it yourself on your own load before you quote it to a client.

The fine print is elsewhere: Ultrafast supports US data residency and global processing only, with no European endpoints. If you hold client data that must not leave the region, this mode is not for you, however fast it is. That is written by OpenAI in its documentation, not by NVIDIA on its blog.

Speed has an address. Here the address is the United States.

What comes next: the price. OpenAI calls it higher and points to a table we are not quoting until we have compared it with Standard on the same task. Whether Ultrafast is for agents in production or for demos is decided by the price of a finished task, not the price of a token.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→NVIDIA - How NVIDIA GPUs Help Accelerate OpenAI's GPT-6 Astra Ultrafast, 01.10.2026→OpenAI - Ultrafast mode, API documentation, read on 02.10.2026
Original: https://wearecoded.com/en/articles/gpt-6-astra-ultrafast-blackwell-8x.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news