Yuanzhi Liang, Xufeng Zhan, et al. · arXiv · 2026
This roadmap places World Action Models inside a broader physical-intelligence stack: an embodied brain compares possible interventions, a physical harness grounds its requests through tools and controllers, shared contracts connect heterogeneous components, and verified interaction becomes post-training experience.
Yuanzhi Liang, Xuan'er Wu, et al. · arXiv · 2026
TeleBoost treats video post-training as a staged, stability-constrained system that connects supervised policy shaping, reward-driven reinforcement learning, and preference-based refinement while diagnosing costly rollouts, compounding temporal failures, and uncertain feedback.
Yuanzhi Liang, Yijie Fang, et al. · Vicinagearth · 2026 · vol. 3(1) · article 2
This survey organizes reinforcement learning for image, video, and 3D/4D generation, treating RL not only as a fine-tuning algorithm but as an interface for optimizing non-differentiable, preference-driven, temporal, and high-level objectives.
Ziqi Ni, Yuanzhi Liang, et al. · IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026) · 2026 · pp. 27260-27269
ViPO turns one scalar reward per generated sample into spatially and temporally structured, pixel-level advantage maps, using pretrained vision features to focus GRPO updates on perceptually important regions while retaining the standard training pipeline.
Rui Li, Bingyu Li, et al. · ACM Multimedia 2026 (accepted) · 2026 · forthcoming
RATS combines trajectory distillation with preference feedback: horizon matching aligns teacher and student at key denoising stages, while a reward-aware gate strengthens teacher guidance only when the teacher is better under the chosen reward.
Rui Li, Yuanzhi Liang, et al. · European Conference on Computer Vision (ECCV 2026, accepted) · 2026 · forthcoming
TaRoS v4 addresses reward hacking and saturation in video GRPO by organizing multi-aspect feedback with component-level performance assessment and intra-group sparsity, then adaptively downweighting components whose scores have saturated.
Rui Li, Ke Hao, et al. · ACM Multimedia 2026 (accepted) · 2026 · forthcoming
OTCA replaces uniform, scalar reward propagation in visual GRPO with structured credit assignment along two axes: which denoising steps matter and which reward objectives should matter at each point in the trajectory.
Ruiying Liu, Yuanzhi Liang, et al. · IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026) · 2026 · pp. 34408-34417
BPGO treats visual-reward scores as uncertain rather than equally trustworthy: a semantic prior anchors inter-group trust allocation and intra-group renormalization so that GRPO emphasizes confident feedback and suppresses ambiguous signals.
Sheng Liu, Yuanzhi Liang, Sidan Du · European Conference on Computer Vision (ECCV 2026, accepted) · 2026 · forthcoming
LaxMotion removes direct 3D pose supervision and instead learns 3D motion as a structurally consistent explanation of global trajectories and monocular 2D kinematic cues, supported by view, orientation, and stability regularization.