Google adds agentic video understanding to Gemini
Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite scan video at a frame rate the model chooses itself, cutting token use by up to 88% on Google's own figures.
- Models & capabilities
- Minor
Google added what it called agentic video understanding to Gemini 3.7 Flash, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, letting the models decide which parts of a video to inspect, at what frame rate, and through which modality — visual frames, audio or transcript — rather than processing every frame at a fixed rate. Google said the approach supports tasks such as sub-second moment retrieval and anomaly detection within long recordings.
The company reported token consumption falling by up to 88% and cost by up to 66% compared with fixed-rate processing, alongside an accuracy improvement of up to 7% on its own evaluations. The feature is available now through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, with a consumer rollout to the Gemini app described as coming soon and YouTube’s “Ask YouTube” feature planned to use it in the following months.
The update targets the cost of processing video at scale — a persistent constraint on using multimodal models over hours of footage — rather than adding a new capability outright, and follows the same variable-effort logic Google and other labs had begun applying to reasoning generally: spend compute where it is needed and skip it where it is not. It arrived days after Google’s video-generation update to Gemini Omni, aimed at the opposite end of the video pipeline — production rather than analysis.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.