Signal Brief

Astribot发布Lumo-2

Astribot发布了40亿参数的Lumo-2机器人基础模型,推理速度比前代快2.71倍,支持人类视频和多机器人身体。该模型在105个未见物体上取得更好结果,并在22项涵盖 motion prediction、memory、physical reasoning、long tasks、fine han...

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-17 收录 2026-07-17 观察

一句话判断

该模型首次将预测与控制融合为潜在世界-动作模型,显著提升推理速度并实现真实家庭任务验证。

核心信息

Astribot发布了40亿参数的Lumo-2机器人基础模型,推理速度比前代快2.71倍,支持人类视频和多机器人身体。该模型在105个未见物体上取得更好结果,并在22项涵盖 motion prediction、memory、physical reasoning、long tasks、fine hand control的真实操作任务中表现最佳,并已在20多个复杂家庭任务的真实机器人上验证。

原始内容

A lot of work being coming out on robot foundation models. Astribot just released Lumo-2, a 4B robot foundation model that learns to predict the real world before deciding how to move. - 2.71x faster total inference, - support for human videos and multiple robot bodies, - stronger results on 105 unseen objects, and - best overall performance across 22 real-world manipulation tasks covering motion prediction, memory, physical reasoning, long tasks, and fine hand control. And these results were validated on a real robot across 20+ complex household tasks, not just benchmark leaderboards. That is a much harder test because the model has to deal with changing scenes, physical contact, timing, memory, and long action sequences where one small mistake can break the entire task. Lumo-2 is a robot control model with a small world model built inside it, I could say its industry's first Latent World-Action Model (LWAM) built for home. In a normal world model, we predict future images or video, and then a separate planner decides what the robot should do. In a normal robot policy, we skip prediction and directly convert the camera view and instructions into motor commands. So Lumo-2 does both those things. First, it only predicts the part of the future that will affect the task, like an object moving, a contact changing or the pouring being finished. Second, it uses that small prediction to move the robot. It doesn’t create an entire video from the future, frame by frame. The model is trained in three steps. First, it links changes of perception to robot motion. Second, it grounds action tokens in language and vision. Third, it is trained on VLM data, videos and robot demonstrations, such that the sequence can anticipate the physical outcome and then act. 🧵 1. https://video.twimg.com/amplify_video/2078168784742002688/vid/avc1/1920x1080/n4sD6Y9aeHnCQ90Y.mp4?tag=29 https://video.twimg.com/amplify_video/2078168793956945920/vid/avc1/1920x1080/q7OApHjuFok-bV-X.mp4?tag=29 https://video.twimg.com/amplify_video/2078170882472849408/vid/avc1/1280x720/X6oKFoo7d0C5w-4s.mp4?tag=29 ![photo](https://pbs.twimg.com/media/HNclbc4bYAAjTIn.jpg) Rohan Paul (@rohanpaul_ai): Technical Report: https://t.co/ECostObGUM Project page: https://t.co/iaM3zyVx2P 📄 Paper: https://t.co/K6C9BtJN97 Rohan Paul (@rohanpaul_ai): 🧵 7. Lumo-2 jumps far beyond Lumo-1 across many embodied reasoning tasks while staying competitive with models built mainly for vision-language benchmarks. The robot training may actually have made its grasp of space and physical scenes more useful. https://t.co/0fa5kWJi3y

相关动态

02

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

BestBlogs 早报 · 07-18

月之暗面发布Kimi K3,2.8万亿参数,896选16的Stable LatentMoE,上下文100万token,接近Fable-5但仍落后最强闭源模型,完整权重7月27日前开源;VentureBeat调查显示54%企业已发生AI代理安全事件;xAI开源Grok Build(84万行Rust代码)并残留上传用户代码痕迹;Cursor评...

twitter关注列表2026-07-17#AI#模型发布#开源
观察