@RohanLikesAI
Today, we’re launching Agentic Video Understanding in Gemini across 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. 📉 ~88% fewer tokens 💰 ~66% lower cost 📈 ~7% higher accuracy The result: a new Pareto frontier for video understanding across accuracy and cost. Traditional video understanding processes video statically at a fixed sampling rate. With Agentic Video Understanding, Gemini can actively decide what to watch, at what speed, and which signals to use (frames, audio, or transcripts) by using native video tools. This enables split-second retrieval, rapid-motion understanding, and needle-in-a-haystack search across hours of video. Available now via the Gemini API - and coming soon to billions of users through the Gemini App and YouTube’s “Ask YouTube” experience. Learn more (details, demos, and docs): https://t.co/kOYrewXHYL PMing this 0→1 has been a career highlight. Taking it from an idea and helping shape it into a frontier-defining capability we’re now shipping broadly has been incredibly rewarding. Huge thanks to @MarioLucic_ , @skprat @ahmetius , @FPavetic, @suhasyogin , @jalayrac, @bcaine, @nbrichtova, @xu_bibo , @tulseedoshi, @davthack, @koraykv and the multimodal team at @GoogleDeepMind.