WHY THIS MATTERS
Google turned on agentic video understanding in Gemini 3.7, 3.6 Flash and 3.5 Flash-Lite. Vendor numbers: up to 88% fewer tokens and 66% lower cost. Not Gemini 3.8.In this article
CONFIRMED: Gemini adds agentic video understanding
Editorial label: CONFIRMED. Basis: Google's own blog on September 1, 2026. Agentic video understanding is live on Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite via the Gemini API in AI Studio and Gemini Enterprise Agent Platform. You set processing to agentic on a video input. No extra feature fee; standard token prices apply. This is a feature on existing Flash models, not Gemini 3.8. Source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-agentic-video-in-gemini/ · docs: https://ai.google.dev/gemini-api/docs/video-understanding#agentic-video-understanding
What we know
Instead of ingesting video at a fixed FPS, the model searches, scans and inspects only the frames, audio and transcript it needs. Google's own evals claim up to 88% fewer tokens, up to 66% lower analysis cost and up to 7% higher accuracy, with the biggest gains on long video. 3.7 Flash sits at their quality-to-cost frontier among the three. The Gemini app rollout is "soon"; YouTube Ask YouTube is months away. Gemini 3.8 Flash is still absent from the official catalog.
What we still do not know
Independent benchmarks are not in. Latency of the agentic loop, Portuguese/Spanish quality and failure modes on noisy user video are untested here. The 88/66/7 figures are vendor numbers. 3.8 Flash remains a WSJ signal, not this launch.
Why it matters
If you already pay Flash to read long YouTube or lecture video, this is the practical drop, not a new model ID. Test processing=agentic on a real long clip before rewriting pipelines. Do not confuse it with a 3.8 Flash ship.