Google's Ask YouTube Gets AI-Powered Video Analysis Upgrade
Google is upgrading its Ask YouTube feature to analyze videos more efficiently and accurately. The new video understanding technology, called agentic video understanding, allows Gemini to focus on specific segments of a video for inspection instead of sampling the entire video at a fixed rate.
The system will power Ask YouTube on the watch page in the coming months, enabling it to answer questions based on what's shown on screen. This feature is already accessible to developers via the Gemini API and Enterprise Agent Platform, supporting both uploaded and YouTube videos.
According to Google, agentic video understanding uses a dynamic loop where the model selectively loads parts of the video, adjusts the frame rate, and decides whether to include frames, audio, or transcripts for specific segments. This method offers several benefits, including locating split-second moments, searching multi-hour videos, spotting visual glitches, and counting repeated actions or objects.
Google reports that in its testing on standard video benchmarks, agentic video understanding reduces token usage by up to 88%, lowers analysis costs by up to 66%, and improves accuracy by up to 7% compared to static processing. The update will change how the watch-page feature can inspect the video you're currently watching.