Real-Time Voice AI Market Heats Up With New Releases from OpenAI, Google's Gemini, and Alibaba's Qwen
The real-time voice AI market has seen significant updates from three major players: OpenAI, Google's Gemini, and Alibaba's Qwen. In a single month, all three companies released new versions of their respective APIs, each with its own pricing structure and latency numbers. The question for developers building voice agents, call-center automation tools, or speech-to-speech translation applications is no longer which chatbot is the smartest, but rather which API to plug into and at what cost per minute.
OpenAI's Realtime API (gpt-realtime-2.1) has been updated with a cheaper variant, gpt-realtime-2.1-mini, cutting audio-input pricing from $32 to $10 per million tokens and audio-output from $64 to $20. Google's Gemini Live API (gemini-3.8-live) is now the default low-latency voice model in the Gemini API, with a maximum session length of around 15 minutes. Alibaba's Qwen team launched Qwen3.8-Omni-Flash on September 14, which undercuts both US rivals by a wide margin, according to TechNode's coverage.
The gap between the cheapest and most expensive option is as high as 213x, with OpenAI's Realtime API being the priciest per million tokens. This significant price difference makes it essential for developers to understand the specifics of each API before committing to one provider.