Gemini's Video Analysis Just Got a Whole Lot Smarter
Google's Gemini AI has become more efficient and accurate at analyzing videos. The latest development brings 'agentic video understanding' to the latest Gemini models, which can now decide what to watch and at what speed. This means Gemini can pinpoint split-second changes, answer complex questions across multi-hour videos, inspect videos for visual artifacts, and even count and track physical movements and objects.
Agentic video understanding is a significant improvement over 'static' processing, which involved splitting the video into individual frames, resulting in slower performance and higher costs. With this new capability, Gemini can reduce token usage by up to 88% and offer up to 7% better accuracy.
The feature is currently available via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, but it will soon roll out to the Gemini app as well. Additionally, Google plans to use agentic video understanding to power YouTube's 'Ask YouTube' feature, which could make it easier for creators to analyze their videos with Gemini.