Signal Brief

Yann LeCun 转发介绍 Patch Policy 架构

Gaoyue Zhou 团队提出 Patch Policy 架构,允许基于 transformer 的策略直接使用密集视觉 token,无需大语言模型,在操纵任务中比微调 7B VLA 性能高 18%,参数量仅为其 0.7%。

twitter关注列表 Yann LeCun (@ylecun) 发布 2026-07-22 收录 2026-07-22 观察

一句话判断

差异点:无需大模型即可实现密集视觉 token 直接用于策略,大幅降低参数量同时提升性能。

核心信息

Gaoyue Zhou 团队提出 Patch Policy 架构,允许基于 transformer 的策略直接使用密集视觉 token,无需大语言模型,在操纵任务中比微调 7B VLA 性能高 18%,参数量仅为其 0.7%。

原始内容

Yann LeCun (@ylecun) 转发了 Gaoyue Zhou (@GaoyueZhou) 的帖子: Pretrained ViTs see the world in rich, dense detail. Most policies pool it to a single vector before acting, discarding most of it. We introduce Patch Policy: a minimal architectural extension that enables transformer-based policies to consume dense tokens directly, no billion-param VLM required. It outperforms a fine-tuned 7B VLA by 18% with ~0.7% of its parameters, enabling robust, precise manipulation. https://video.twimg.com/amplify_video/2080012982998687744/vid/avc1/2230x1080/-zROs41NY2ZEOdho.mp4?tag=29

相关动态

04

利用 Embedding 模型实现图像打标

Han Xiao 通过将冻结的 jina-v5-omni 多模态嵌入模型进行测试时缩放(scaled at test time),在无需训练、无需第二模型及外部知识的硬约束下,实现了强大的开放词汇多标签 n-gram 图像打标器。

twitter关注列表2026-07-22#技术#研究#模型
观察
05

AI 30年图论猜想被推翻

Dmitry Rybin 通过 GPT 5.6 Pro 发现 Dinitz-Garg-Goemans 猜想错误,该猜想开放约30年。具体反例图显示分式流成本58,不可分裂流(容量违反≤15)成本至少60。

twitter关注列表2026-07-22#AI#技术突破#研究
观察