Gemini AI Model Upgraded with Agent-Based Video Analysis
Google has upgraded its Gemini AI model with agent-based video analysis, enabling it to scan videos more efficiently and accurately. This new approach allows the model to hunt for relevant sections on its own, cutting token usage by up to 88 percent.
The latest models, including Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, can now pick up moments shorter than one second, making automated video editing more precise. The system also tracks individual scenes in hours of footage without burning through millions of tokens.
Google says the agent-based variant ties the model's reasoning directly to native video tools, deciding on its own which sections to look at and how to process them. This approach builds on 'agentic vision' introduced with Gemini 3 Flash in January.
The efficiency gains are most noticeable with long videos, where static processing often led to high token costs or missed details. On Google's benchmarks, including LongVideoBench, Gemini 3.7 Flash with agent-based analysis delivers the best mix of accuracy and cost efficiency.