论文解读 · 作者已确认

TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model

作者: Yabo Chen, , Jiepeng Wang, Tingxi Chen, Junfei Cheng, Zixiao Gu, Yuyang Huang, Zicheng Jiang, Wei Li, Tian Li, Weichen Li, Zuoxin Li, Guangce Liu, Jialun Liu, Junqi Liu, Haoyuan Wang, Qizhen Weng, Xuan'er Wu, Xunzhi Xiang, Xiaoyan Yang, Xin Zhang, Shiwen Zhang, Junyu Zhou, Chengcheng Zhou, Haibin Huang, Chi Zhang, Xuelong Li

arXiv · 2025

4D world modelvideo generationdynamic reconstructionlong-term memoryreal-time synthesis

论文信息

作者
Yabo Chen, Yuanzhi Liang, Jiepeng Wang, Tingxi Chen, Junfei Cheng, Zixiao Gu, Yuyang Huang, Zicheng Jiang, Wei Li, Tian Li, Weichen Li, Zuoxin Li, Guangce Liu, Jialun Liu, Junqi Liu, Haoyuan Wang, Qizhen Weng, Xuan'er Wu, Xunzhi Xiang, Xiaoyan Yang, Xin Zhang, Shiwen Zhang, Junyu Zhou, Chengcheng Zhou, Haibin Huang, Chi Zhang, and Xuelong Li
推荐论文引用
Yabo Chen, Yuanzhi Liang, Jiepeng Wang, Tingxi Chen, Junfei Cheng, Zixiao Gu, Yuyang Huang, Zicheng Jiang, Wei Li, Tian Li, Weichen Li, Zuoxin Li, Guangce Liu, Jialun Liu, Junqi Liu, Haoyuan Wang, Qizhen Weng, Xuan'er Wu, Xunzhi Xiang, Xiaoyan Yang, Xin Zhang, Shiwen Zhang, Junyu Zhou, Chengcheng Zhou, Haibin Huang, Chi Zhang, and Xuelong Li. “TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model.” arXiv (2025). arXiv:2601.00051.

版本日期

arXiv 首次提交
2025-12-31
arXiv 最近修订
2025-12-31
来源核验日期
2026-07-31
核验状态
作者已于 2026-07-31 核验

论文概要

TeleWorld 在视频生成和 4D 重建之间形成闭环:生成流持续写入时空表示,这个持久状态再指导后续生成;Macro-from-Micro Planning 与蒸馏用于支持长时和低延迟合成。

研究问题

高质量视频生成器仍缺少实时交互、动态场景的持久记忆,以及长时间范围内可靠的空间、时间和物理一致性,因此尚不能直接作为实用 world model。

论文贡献

  • 在闭环 4D 框架中统一视频生成、动态场景重建与长期世界记忆。
  • 提出 generation–reconstruction–guidance 循环,用重建后的 4D 状态指导下一段生成。
  • 结合 Macro-from-Micro Planning 与 Distribution Matching Distillation,面向长时低延迟合成。

证据与评测范围

报告评测了静态与动态世界理解、长时一致性和实时生成效率;准确延迟、硬件、场景与 benchmark 条件应以实验章节为准。

适用范围与局限

闭环系统可能同时累积重建误差与生成误差,长时表现取决于状态被重建和复用的准确性;“实时”结论也受论文所用算力和配置约束。

Related work 定位

TeleWorld 连接视频生成、动态重建和持久记忆,适合被定位为从一次性视频生成走向有状态、可交互 world model 的工作。

Related Work 表述

Chen 等提出 TeleWorld:将生成视频重建为持久 4D 时空状态并用于指导后续生成,同时以层次规划和蒸馏支持长时实时合成。

这是一段用于说明论文定位的简洁中性表述。

论文官方英文摘要

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual quality, they remain limited in real-time interaction, long-horizon consistency, and persistent memory of dynamic scenes, hindering their evolution into practical world models. In this report, we present TeleWorld, a real-time multimodal 4D world modeling framework that unifies video generation, dynamic scene reconstruction, and long-term world memory within a closed-loop system. TeleWorld introduces a novel generation-reconstruction-guidance paradigm, where generated video streams are continuously reconstructed into a dynamic 4D spatio-temporal representation, which in turn guides subsequent generation to maintain spatial, temporal, and physical consistency. To support long-horizon generation with low latency, we employ an autoregressive diffusion-based video model enhanced with Macro-from-Micro Planning (MMPL)--a hierarchical planning method that reduces error accumulation from frame-level to segment-level-alongside efficient Distribution Matching Distillation (DMD), enabling real-time synthesis under practical computational budgets. Our approach achieves seamless integration of dynamic object modeling and static scene representation within a unified 4D framework, advancing world models toward practical, interactive, and computationally accessible systems. Extensive experiments demonstrate that TeleWorld achieves strong performance in both static and dynamic world understanding, long-term consistency, and real-time generation efficiency, positioning it as a practical step toward interactive, memory-enabled world models for multimodal generation and embodied intelligence.

摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。

依据与出处

核对内容论文中的位置
问题陈述Abstract
方法与贡献Abstract; generation-reconstruction-guidance, MMPL, and DMD sections
评测结论Abstract; static/dynamic understanding, consistency, and efficiency experiments

主要核验来源: arXiv record (arXiv:2601.00051v1).

如何引用

科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。

Yabo Chen, Yuanzhi Liang, Jiepeng Wang, Tingxi Chen, Junfei Cheng, Zixiao Gu, Yuyang Huang, Zicheng Jiang, Wei Li, Tian Li, Weichen Li, Zuoxin Li, Guangce Liu, Jialun Liu, Junqi Liu, Haoyuan Wang, Qizhen Weng, Xuan'er Wu, Xunzhi Xiang, Xiaoyan Yang, Xin Zhang, Shiwen Zhang, Junyu Zhou, Chengcheng Zhou, Haibin Huang, Chi Zhang, and Xuelong Li. “TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model.” arXiv (2025). arXiv:2601.00051.

复用许可

本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.

一手来源与资源