论文概要
这篇路线图把 World Action Models 放入更完整的物理智能系统:embodied brain 比较候选干预,physical harness 通过工具和控制器将请求落地,共享 contracts 连接异构组件,经过验证的交互再转化为后训练经验。
研究问题
Action model、VLA policy 与 world model 虽然持续发展,但它们在表示、目标、数据集、任务约定和运行时接口上彼此割裂,难以被组合、复用、统一评测并形成持续改进的系统。
论文贡献
- 把领域瓶颈归纳为模型角色与表示、目标与标准化、系统组合三个相互耦合的缺口。
- 提出 embodied brain 这一长期目标:结合多模态上下文比较候选干预,输出状态转移或能力请求,而不是直接下发执行器指令。
- 以 physical harness、共享 contracts 和闭环后训练作为模型落地、组件连接、行为验证与经验复用的系统机制。
证据与评测范围
本文属于综述与路线图,证据主要来自对 action model、VLA 和 world model 文献的结构化梳理以及系统设计论证;它并未声称已经实现一个完整部署的 embodied brain 或新的统一基准。
适用范围与局限
该方案仍是面向未来的研究议程。不同 embodiment 下的接口规范、验证机制、harness 实现与闭环训练流程都需要进一步工程化和实证检验。
Related work 定位
在讨论预测型 world/action model 如何成为开放世界具身系统的可复用组件时可以引用这项工作。其差异点是把模型角色、标准接口、运行时落地和验证交互驱动的学习放在同一条演进路线中。
Related Work 表述
Liang 等提出了从 World Action Models 走向 embodied brains 的路线图,主张将预测模型与 physical harness、共享 contracts 和闭环后训练结合,以构建模块化的开放世界物理智能系统。
这是一段用于说明论文定位的简洁中性表述。
论文官方英文摘要
Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect candidate interventions with predicted consequences. However, progress remains fragmented: models use incompatible action spaces and prediction targets, datasets and tasks follow different conventions, and runtime systems expose limited interfaces for reuse and evaluation. We review the evolution toward WAMs and organize these limitations into three coupled gaps: model roles and representations, objectives and standardization, and system composition. Building on this analysis, we propose a co-evolution roadmap for physical intelligence centered on the embodied brain, a long-term model target for integrating multimodal context, comparing candidate interventions, and issuing state-transition or capability requests rather than direct actuator commands. WAMs provide promising prototypes for its predictive functions, while a physical harness grounds model outputs through tools, controllers, verification, and trace logging. Shared contracts align heterogeneous models, data, tasks, and embodiments, and closed-loop post-training converts verified interaction into reusable experience. Together, these components define a modular physical-intelligence stack for adaptive and self-improving embodied agents.
摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。
依据与出处
| 核对内容 | 论文中的位置 |
|---|---|
| 问题陈述 | Abstract |
| 方法与贡献 | Abstract; roadmap sections on embodied brains and physical harnesses |
| 评测结论 | Abstract; paper synthesis and roadmap discussion |
主要核验来源: arXiv record (arXiv:2607.11689v1).
如何引用
科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。
Yuanzhi Liang, Xufeng Zhan, Haibin Huang, Chi Zhang, and Xuelong Li. “From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence.” arXiv (2026). arXiv:2607.11689.
复用许可
本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.