Signal Brief

FLUX 3 将极大影响机器人智能和学习

Black Forest Labs 发布 FLUX 3 统一多模态模型,支持图像、视频、音频和动作预测。其 Self-Flow 架构使机器人微调数据减少一半,相比早期模仿视频实验,机器人训练数据减少 10 倍,通过大规模视频预训练隐式学习物理规律。

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-23 收录 2026-07-23 观察

一句话判断

差异点:FLUX 3 的 Self-Flow 架构提供内部分布预测,大幅降低机器人对真实演示数据的依赖,与标准 VLA 的静态预训练+少量演示范式形成关键区别。

核心信息

Black Forest Labs 发布 FLUX 3 统一多模态模型,支持图像、视频、音频和动作预测。其 Self-Flow 架构使机器人微调数据减少一半,相比早期模仿视频实验,机器人训练数据减少 10 倍,通过大规模视频预训练隐式学习物理规律。

原始内容

FLUX 3 will have big positive impact on robot intelligence and learning. Standard vision-language-action models learn their meanings via pretraining on static images and text, and then need to learn about contact, deformation, temporal causality, and recovery from a few teleoperated robot paths. Because static image–text pretraining does not teach how objects behave over time, standard VLAs must learn physics from limited robot demonstrations. But FLUX 3 brings those motion and interaction patterns from large-scale video pretraining. The FLUX 3 backbone is trained to generate temporally consistent video, jointly with images and audio. So to predict future frames, it must encode object motion, contact transitions, deformation, persistence and event ordering. FLUX-mimic does not directly guess robot movements from the current camera view. FLUX 3 first builds an internal picture of what should happen next—how the hand, object and scene are likely to move. A smaller control model then converts that predicted motion into joint and gripper commands. The future video is not actually generated, so the system can still react quickly. The important architectural detail is Self-Flow. Many video models can produce realistic footage without organizing their internal knowledge in a way a robot can easily use. Self-Flow trains FLUX 3 to share information more clearly across frames, objects and modalities. BFL says this lets the robot learn tasks with roughly half the fine-tuning, while earlier mimic-video experiments used up to 10× less robot training data than a comparable VLA. https://video.twimg.com/amplify_video/2080416170122080256/vid/avc1/1280x720/NNYZHaPTXqiMmmlN.mp4?tag=29 > **引用原帖 Black Forest Labs (@bfl_ai):** > Introducing FLUX 3. > One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. > FLUX 3 Video is now available in early access (link below). > Jointly trained in one unified architecture, our model can be extended to predict actions for robotics. See our work with mimic and Audi in the thread. > https://x.com/bfl_ai/status/2080308988961554582

相关动态

03

Black Forest Labs 发布 FLUX 3 统一多模态模型

Black Forest Labs 发布 FLUX 3,采用统一架构同时处理图像生成、视频、原生音频和机器人动作预测。视频已开放早期访问,后续将开放权重,并与 mimic、Audi 合作在真实机器人上运行。团队认为物理世界智能与内容创作可共享同一视觉基础模型。

twitter关注列表2026-07-23#模型发布#多模态#技术突破
观察