Qwen's Omni-Flash Outperforms Gemini at a Fraction of the Cost
Qwen3.8-Omni-Flash is Qwen's first multimodal model designed for AI agents that can process audio and video together, making conclusions based on this information.
The model can be used to edit vlogs, translate short videos, or summarize movies using its built-in tools.
Qwen3.8-Omni-Flash has a context window of one million tokens and matches the multimodal benchmarks of Google's Gemini 3.8 Flash in audio-video tasks.
API pricing for Qwen3.8-Omni-Flash starts at $0.15 per million input tokens and $0.47 per million output tokens, significantly undercutting Gemini's introductory rate of $0.75 for input and $3.75 for output per million tokens.