Signal Brief

Google Gemini Managed Agents 更新

Google 更新 Gemini API Managed Agents,新增后台任务、远程 MCP 支持、函数调用、凭据刷新和免费层访问,使代理更接近生产环境。后台任务允许长时间运行的操作,远程 MCP 可直接访问私有服务,无需自定义代理。

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-07 收录 2026-07-07 观察

一句话判断

新增的后台任务和远程MCP功能解决了生产环境中常见的阻塞点,值得关注具体实现细节。

核心信息

Google 更新 Gemini API Managed Agents,新增后台任务、远程 MCP 支持、函数调用、凭据刷新和免费层访问,使代理更接近生产环境。后台任务允许长时间运行的操作,远程 MCP 可直接访问私有服务,无需自定义代理。

原始内容

Google Gemini just gave developers much stronger tools for production AI agents. i.e. Gemini API Managed Agents now are much closer to production with new addition of background tasks, remote MCP, function calls, credential refresh, and free-tier access. Managed Agents are Google-hosted AI workers that run antigravity-preview-05-2026 inside an isolated Linux sandbox. Older agent apps often broke when a long task outlived a normal HTTP request. An "Interaction" is the important object here. It stores the task, the model’s steps, tool calls, tool results, and final output. So instead of your app manually tracking every turn, tool result, and file, Google tracks much of that server-side. Remote MCP support changes the tool story, because agents can contact private services without custom proxy glue. A company can now connect observability, databases, or internal APIs beside Google Search and code execution. Function calling adds another split, where Google runs sandbox tools and your app handles business logic. Credential refresh fixes a production pain, since short-lived tokens can rotate without losing sandbox state. Overall, this makes Gemini API feel less like a model endpoint and more like agent infrastructure. ![photo](https://pbs.twimg.com/media/HMpQqv8bgAAXwPv.jpg) > **引用原帖 Google AI Studio (@GoogleAIStudio):** > https://t.co/ohyGvAO3w8 > https://x.com/GoogleAIStudio/status/2074533418004591077

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
03

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

开源模型长期网络能力差距缩小至4-7个月

长期网络能力方面,领先开源模型落后闭源前沿从2025年大部分时间的6-10个月缩短至4-7个月。GLM-5.2匹配了约4个月前发布的闭源模型。AI安全研究所的32步“The Last Ones”任务中,GPT-5.6 Sol在10次尝试中完成7次,每次预算1亿token,且性能随推理token增加而提升。

twitter关注列表2026-07-17#模型#技术突破#AI安全
观察
05

Runway Agent 在第三方评测中全面领先

Physion Labs 发布 Physion-Arc 1.0 基准测试,对 Runway、Luma、MiniMax、Kling、Utopia 和 TapNow 六款 AI 视频代理进行独立人类评估,使用 30 个电影提示和 16 个指标。Runway Agent 2.0 在叙事连贯性、电影语言和制作质量三项核心维度全部排名第一。

twitter关注列表2026-07-17#评测#多模态#模型
观察