Signal Brief

Runway Agent 在第三方评测中全面领先

Physion Labs 发布 Physion-Arc 1.0 基准测试,对 Runway、Luma、MiniMax、Kling、Utopia 和 TapNow 六款 AI 视频代理进行独立人类评估,使用 30 个电影提示和 16 个指标。Runway Agent 2.0 在叙事连贯性、电影语言和制...

twitter关注列表 Runway (@runwayml) 发布 2026-07-17 收录 2026-07-17 观察

一句话判断

值得点开原文查看详细评测结果,了解 Runway 在具体指标上的优势。

核心信息

Physion Labs 发布 Physion-Arc 1.0 基准测试,对 Runway、Luma、MiniMax、Kling、Utopia 和 TapNow 六款 AI 视频代理进行独立人类评估,使用 30 个电影提示和 16 个指标。Runway Agent 2.0 在叙事连贯性、电影语言和制作质量三项核心维度全部排名第一。

原始内容

Runway Agent ranked No. 1 across every dimension. An independent human evaluation put six AI video agents through 30 cinematic prompts and 16 metrics. Runway came out on top in all three pillars: Narrative Coherence, Cinematic Language, and Production Quality. https://t.co/aBwqH3hUuL ![photo](https://pbs.twimg.com/media/HNdOSvHbQAAu92o.jpg) > **引用原帖 Physion Labs Official (@Physion_Labs):** > 🎬 Video agents can now generate minute-long videos. But can they actually direct? > Today, we’re launching 𝐏𝐡𝐲𝐬𝐢𝐨𝐧-𝐀𝐫𝐜 1.0, a new benchmark evaluating complete, multi-scene videos across narrative coherence, cinematic language, and production quality. > We tested @runwayml, @LumaLabsAI, @MiniMax_AI, @Kling_ai , @UtopaiStudios and @TapNow_AI on 100 screenplays and 600 generated videos. > 🏆 𝐑𝐮𝐧𝐰𝐚𝐲 𝐀𝐠𝐞𝐧𝐭 2.0 𝐫𝐚𝐧𝐤𝐞𝐝 𝐍𝐨. 1 𝐨𝐯𝐞𝐫𝐚𝐥𝐥 𝐚𝐧𝐝 𝐥𝐞𝐝 𝐞𝐯𝐞𝐫𝐲 𝐞𝐯𝐚𝐥𝐮𝐚𝐭𝐢𝐨𝐧 𝐝𝐢𝐦𝐞𝐧𝐬𝐢𝐨𝐧. Its advantage was especially clear across subjective metrics, where cinematic taste matters most. Runway ranked first on all eight. > 🔗 Read the full benchmark: https://t.co/PuBpmwUQGv > https://x.com/Physion_Labs/status/2077551810357969047 Runway (@runwayml): Learn more about our approach to building Runway Agent: https://t.co/V8vctepG4Z

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
03

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

开源模型长期网络能力差距缩小至4-7个月

长期网络能力方面,领先开源模型落后闭源前沿从2025年大部分时间的6-10个月缩短至4-7个月。GLM-5.2匹配了约4个月前发布的闭源模型。AI安全研究所的32步“The Last Ones”任务中,GPT-5.6 Sol在10次尝试中完成7次,每次预算1亿token,且性能随推理token增加而提升。

twitter关注列表2026-07-17#模型#技术突破#AI安全
观察
05

Kimi K3 编码代理评测

Artificial Analysis发布Kimi K3在编码代理指数上的评测结果:得分57,排名第5,性能与GPT-5.6 Terra和GPT-5.5持平(57),超过Opus 4.8(55),略低于Grok 4.5(58)和Fable 5(59);成本平均3.18美元/任务,比GPT-5.6 Sol便宜55%。

twitter关注列表2026-07-17#AI模型#评测
观察