Google Cuts Video Analysis Costs by Up to 88% with Agentic Video Understanding
Google has announced the launch of agentic video understanding across its Flash models. This innovation allows Gemini models to navigate video timelines instead of ingesting them at a fixed one frame per second.
This approach enables the model to decide what to watch, at what frame rate, and through which modality, resulting in significant efficiency gains. Google reports that agentic processing leads to up to 88% fewer tokens, 66% lower cost, and 7% higher accuracy on standard video benchmarks.
The new feature is currently only deployable as a hosted API feature, with no open weights or self-hosting options available. However, it can be enabled through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.