Agentic Video Understanding in Gemini
Agentic video analysis for faster, smarter Gemini insights
评分 82 ▲ 137 ◆ 1 付费 PH 精选
概述
Google Gemini 推出智能体视频理解模式,让模型自主决定观看策略,大幅降低 token 消耗和成本,同时提升长视频分析准确率。
创新点
模型自主决定视频观看速度、帧率和模态(而非固定帧率处理),实现自适应视频理解,token 消耗降低最高 88%,成本降低最高 66%,是视频 AI 处理范式的重要转变
目标用户
使用 Gemini API 的开发者、企业 AI 应用构建者、需要处理长视频内容的 AI 产品团队
竞争格局
竞争对手包括 OpenAI GPT-4o 视频能力、Anthropic Claude 多模态、AWS Rekognition 等,但 Google 凭借 Gemini 原生多模态优势和成本优化具备差异化竞争力
标签
API人工智能视频分析开发者工具Google/Gemini多模态AI
团队
- Rohan Doshi — ML @ Waymo | xGoogle, Princeton CS
相关产品
GPT-6 Astra 94
OpenAI's most capable model for end-to-end work
Claude Fable 5 91
Anthropic's most capable model ever — free until June 22
AlphaGenome Atlas 88
Google's AI map of every possible human DNA mutation
WeatherNext 3 88
Our most advanced and accurate global weather AI model
Gemini Robotics 2 88
Google's AI brain for the next generation of robots
GPT-5.6 88
A new standard for intelligence and efficiency