Google Transforms Gemini Video Analysis with Agent-Based Approach
Google has added agent-based video analysis capabilities to three versions of its AI model 'Gemini' (3.7 Flash, 3.6 Flash, 3.5 Flash-Lite). In the new approach, the model autonomously determines which scenes to analyze and at what resolution, reducing token usage by up to 88% compared to traditional uniform frame loading, according to Google. The company also expects improved accuracy, with particularly significant benefits for processing long-duration videos.

Google has added new video analysis capabilities to multiple versions of its AI model 'Gemini'. The targeted models are Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, which will transition to an 'agent-based' approach that fundamentally differs from the traditional processing method.
Previously, video analysis typically relied on frame-by-frame loading at fixed intervals, similar to stepping through individual frames. While this approach is straightforward, it consumes substantial processing resources even for scenes with minimal movement. Additionally, processing costs—measured in 'tokens,' the computational unit used in AI—scale proportionally with video length, making it particularly inefficient for handling multi-hour video content.
With the new agent-based approach, the model itself decides which scenes to prioritize and analyze. Rather than processing the entire video uniformly, it selects and analyzes portions deemed important. Furthermore, the model automatically determines the resolution level at which to process each scene, enabling detailed analysis where needed and lighter processing elsewhere.
According to Google, this transition can reduce token usage by up to 88%. Simultaneously, the company expects improved analysis accuracy, with particularly pronounced benefits for long-duration videos. The reduction in token consumption directly translates to lower processing costs, potentially resulting in significant cost savings for developers and enterprises using Gemini via API.
This development aligns with intensifying competition over AI model processing efficiency. In business applications handling large-scale video data, cost efficiency is becoming as critical as accuracy in model selection. Demand continues to rise for systems that automatically analyze extended recordings, making cost-effectiveness paired with high accuracy a key differentiator in model competitiveness.
Agent-based analysis is grounded in the concept of models making autonomous decisions while processing tasks, applying the 'AI agent' paradigm—increasingly prevalent in the AI field—to internal model operations. This approach, which adapts decision-making based on context rather than following rigid rules, is recognized as central to AI performance improvements across text processing and search applications. Full-scale implementation in video analysis represents a significant turning point.
Future focus will be on determining the extent to which this approach is adopted in practical applications and whether similar methodologies emerge across competing models and services. For organizations considering deployment of video processing systems and surveillance-analysis tools, accurately evaluating real-world cost reduction benefits and accuracy improvements will be critical to their next decision-making phase.
This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.