OpenAI Tunes Astra Ultrafast on Nvidia GPUs with Internal Models
OpenAI has optimized its Astra Ultrafast service on Nvidia GPUs using its own models. In a post published by Nvidia, Uday Ruddarraju, chief technology officer of compute at OpenAI, said: 'We used our internal models to optimize inference on NVIDIA GPUs... and NVIDIA's programmability helped us deliver the acceleration behind Astra Ultrafast.'
The service is part of GPT-6 Astra and runs on Nvidia Blackwell GPUs. Ultrafast offers up to 8x faster token generation than the standard mode, according to Nvidia. However, this figure refers only to token generation time, not the overall request processing time.
Nvidia's own documentation notes that the 'up to' label is used for vendor figures, and no benchmark or configuration details are provided. OpenAI's API changelog entry states that Ultrafast exists to 'reduce the time between generated output tokens.'
The service comes with a price tag: $300 per million output tokens, six times the standard rate.