Paper overview · author-verified

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

Authors: , Xufeng Zhan, Haibin Huang, Chi Zhang, Xuelong Li

arXiv · 2026

world action modelsembodied intelligencephysical intelligenceworld modelsrobotics

Publication details

Authors
Yuanzhi Liang, Xufeng Zhan, Haibin Huang, Chi Zhang, and Xuelong Li
Recommended paper citation
Yuanzhi Liang, Xufeng Zhan, Haibin Huang, Chi Zhang, and Xuelong Li. “From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence.” arXiv (2026). arXiv:2607.11689.

Version dates

arXiv first posted
2026-07-13
arXiv last revised
2026-07-13
Source checked
2026-07-31
Verification status
Author-verified on 2026-07-31

Summary

This roadmap places World Action Models inside a broader physical-intelligence stack: an embodied brain compares possible interventions, a physical harness grounds its requests through tools and controllers, shared contracts connect heterogeneous components, and verified interaction becomes post-training experience.

Research paths

How this paper contributes to the site's broader research map.

Research question

Research on action models, vision-language-action policies, and world models is advancing, but incompatible representations, objectives, datasets, tasks, and runtime interfaces make the resulting systems difficult to compose, evaluate, and improve as a whole.

What the paper contributes

  • Organizes the field's limitations into coupled gaps in model roles and representations, objectives and standardization, and system composition.
  • Defines the embodied brain as a long-term model target that reasons over multimodal context and requests state transitions or capabilities instead of directly commanding actuators.
  • Proposes physical harnesses, shared contracts, and closed-loop post-training as the system mechanisms that ground, connect, verify, and reuse model behavior.

Evidence and evaluation scope

The paper is a review and roadmap. Its support is a structured synthesis of prior action-model, VLA, and world-model research and a systems argument for the proposed stack; it does not present a newly deployed embodied system or a standalone empirical benchmark.

Scope and limitations

The architecture is a forward-looking research agenda. Individual contracts, verification mechanisms, harness implementations, and closed-loop training procedures still require concrete specifications and empirical validation across embodiments.

Positioning for related work

Use this work when discussing how predictive world/action models can become reusable components of open-world embodied systems. Its distinctive contribution is the co-design of model roles, standardized interfaces, runtime grounding, and learning from verified interaction.

Related-work context

Liang et al. present a roadmap from World Action Models to embodied brains, arguing that predictive models should be integrated with physical harnesses, shared contracts, and closed-loop post-training to support modular open-world physical intelligence.

A concise, neutral description of how this paper can be situated in related work.

Official abstract

Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect candidate interventions with predicted consequences. However, progress remains fragmented: models use incompatible action spaces and prediction targets, datasets and tasks follow different conventions, and runtime systems expose limited interfaces for reuse and evaluation. We review the evolution toward WAMs and organize these limitations into three coupled gaps: model roles and representations, objectives and standardization, and system composition. Building on this analysis, we propose a co-evolution roadmap for physical intelligence centered on the embodied brain, a long-term model target for integrating multimodal context, comparing candidate interventions, and issuing state-transition or capability requests rather than direct actuator commands. WAMs provide promising prototypes for its predictive functions, while a physical harness grounds model outputs through tools, controllers, verification, and trace logging. Shared contracts align heterogeneous models, data, tasks, and embodiments, and closed-loop post-training converts verified interaction into reusable experience. Together, these components define a modular physical-intelligence stack for adaptive and self-improving embodied agents.

The abstract is reproduced for scholarly identification and remains under the paper publisher/authors’ original copyright; it is not covered by this page’s CC BY license.

Evidence references

What to verifyLocation in the paper
Problem statementAbstract
Method and contributionsAbstract; roadmap sections on embodied brains and physical harnesses
Evaluation statementAbstract; paper synthesis and roadmap discussion

Primary source: arXiv record (arXiv:2607.11689v1).

How to cite

Cite the paper—not this explainer—for scientific claims. Cite this page only when reusing its original commentary.

Yuanzhi Liang, Xufeng Zhan, Haibin Huang, Chi Zhang, and Xuelong Li. “From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence.” arXiv (2026). arXiv:2607.11689.

Reuse policy

Original explanatory text on this page is licensed under CC BY 4.0 with attribution and a link to this page. Paper title, abstract, figures, and bibliographic metadata are excluded and retain their original rights. Creative Commons Attribution 4.0 International.

Primary sources and resources