Google adds video agent to Gemini API, with analysis costs cut by as much as 66%

Google adds video agent to Gemini API, with analysis costs cut by as much as 66%

N
News Editor
2026-09-02 04:03:36
Google has rolled out Agentic Video Understanding for the Gemini API, adding a mode that lets the model decide which parts of a video to inspect and how to inspect them. In the standard setup, video is sampled at a fixed 1 frame per second and then loaded into context all at once. The new mode changes that workflow: Gemini can search along the timeline for relevant segments and choose whether to use visual frames, audio, or transcript text based on the question being asked. If it encounters fast motion, it can raise the frame rate and check again. According to Google’s own testing, enabling the feature on Gemini 3.7 Flash reduced token consumption by as much as 88%, lowered analysis costs by as much as 66%, and improved relative accuracy by up to about 7%. Google said the system can identify moments shorter than one second and also find specific details buried inside videos that run for hours. The feature is now available on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, supports both uploaded videos and YouTube videos, and does not carry an extra feature charge. Google also said it plans to bring the capability to the Gemini App and YouTube’s Ask YouTube.

Google has launched Agentic Video Understanding in the Gemini API, giving the model the ability to decide which parts of a video to inspect and how to check them.

In the standard mode, the system samples video at a fixed 1 frame per second and places the extracted frames into context in one batch. The new mode works differently. It can search along the timeline for relevant clips and choose visual frames, audio, or transcript text depending on the query. When it runs into fast-moving action, it can increase the frame rate and inspect the segment again.

In Google’s internal testing, enabling the feature on Gemini 3.7 Flash cut token usage by as much as 88%, reduced analysis costs by as much as 66%, and improved relative accuracy by up to about 7%.

Google said the system can pinpoint moments shorter than one second and can also find a specific detail inside videos that last for hours. From a technical standpoint, the feature brings into Gemini an internal version of the workflow developers previously had to build themselves: first locate the relevant segment, then examine it in greater detail.

The feature is now supported on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. It works with both uploaded videos and YouTube videos, and Google said there is no separate charge for the capability. The company added that it will later integrate the feature into the Gemini App and YouTube’s Ask YouTube.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
8000

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.