Publications

Browse publications by year. Each paper page includes the abstract, a concise overview, related-work context, sources, and citation downloads.

Explore by research path

中文目录

Original commentary is CC BY 4.0; paper abstracts and bibliographic material retain their original rights.

2026 9

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

Yuanzhi Liang, Xufeng Zhan, et al. · arXiv · 2026

This roadmap places World Action Models inside a broader physical-intelligence stack: an embodied brain compares possible interventions, a physical harness grounds its requests through tools and controllers, shared contracts connect heterogeneous components, and verified interaction becomes post-training experience.

TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation

Yuanzhi Liang, Xuan'er Wu, et al. · arXiv · 2026

TeleBoost treats video post-training as a staged, stability-constrained system that connects supervised policy shaping, reward-driven reinforcement learning, and preference-based refinement while diagnosing costly rollouts, compounding temporal failures, and uncertain feedback.

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation

Ziqi Ni, Yuanzhi Liang, et al. · IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026) · 2026 · pp. 27260-27269

ViPO turns one scalar reward per generated sample into spatially and temporally structured, pixel-level advantage maps, using pretrained vision features to focus GRPO updates on perceptually important regions while retaining the standard training pipeline.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Rui Li, Yuanzhi Liang, et al. · European Conference on Computer Vision (ECCV 2026, accepted) · 2026 · forthcoming

TaRoS v4 addresses reward hacking and saturation in video GRPO by organizing multi-aspect feedback with component-level performance assessment and intra-group sparsity, then adaptively downweighting components whose scores have saturated.

Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation

Ruiying Liu, Yuanzhi Liang, et al. · IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026) · 2026 · pp. 34408-34417

BPGO treats visual-reward scores as uncertain rather than equally trustworthy: a semantic prior anchors inter-group trust allocation and intra-group renormalization so that GRPO emphasizes confident feedback and suppresses ambiguous signals.

2025 3

Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts

Sheng Liu, Yuanzhi Liang, et al. · SIGGRAPH Asia 2025 Conference Papers · 2025 · pp. 1-11

Uni-Inter represents human–human, human–object, and human–scene interactions in one Unified Interactive Volume and predicts motion probabilistically joint by joint, enabling one task-agnostic model to reason over heterogeneous and compound interaction contexts.

InterSyn: Interleaved Learning for Dynamic Motion Synthesis in the Wild

Yiyi Ma, Yuanzhi Liang, et al. · IEEE/CVF International Conference on Computer Vision (ICCV 2025) · 2025 · pp. 12832-12841

InterSyn learns solo and multi-person dynamics together rather than treating them as separate tasks: INS synthesizes both from a first-person interaction perspective, and REC refines relative coordination so characters move synchronously.

2024 5

Penalizing the Hard Example But Not Too Much: A Strong Baseline for Fine-Grained Visual Classification

Yuanzhi Liang, Linchao Zhu, et al. · IEEE Transactions on Neural Networks and Learning Systems · 2024 · vol. 35(5) · pp. 7048-7059

MHEM means Moderate Hard Example Modulation: it defines conditions for a loss that emphasizes informative hard samples without continually increasing the influence of extreme cases that a fine-grained classifier may simply memorize.

IcoCap: Improving Video Captioning by Compounding Images

Yuanzhi Liang, Linchao Zhu, et al. · IEEE Transactions on Multimedia · 2024 · vol. 26 · pp. 4389-4400

IcoCap changes the content density seen by a video captioner: Image-Video Compounding Strategy (ICS) injects concise image semantics into video samples, while Visual-Semantic Guided Captioning (VGC) adapts caption supervision to the resulting ambiguous compound content.

FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention

Yu Lu, Yuanzhi Liang, et al. · Advances in Neural Information Processing Systems (NeurIPS 2024) · 2024 · vol. 37 · pp. 131434-131455

FreeLong extends pretrained short-video diffusion models without additional training by blending low-frequency global features with high-frequency local-subsequence features during denoising, targeting global consistency and local spatiotemporal detail at once.

2023 1

MAAL: Multimodality-Aware Autoencoder-based Affordance Learning for 3D Articulated Objects

Yuanzhi Liang, Xiaohan Wang, et al. · IEEE/CVF International Conference on Computer Vision (ICCV 2023) · 2023 · pp. 217-227

MAAL learns affordances for 3D articulated objects with a one-stage autoencoder-based pipeline and a MultiModal Energized Encoder that jointly models object geometry, robot actions, and their interactions while requiring only a small number of positive samples.

2022 2

SEEG: Semantic Energized Co-speech Gesture Generation

Yuanzhi Liang, Qianyu Feng, et al. · IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022) · 2022 · pp. 10463-10472

SEEG separates beat-related and semantic information with a DEcoupled Mining module, then uses a Semantic Energizing Module and semantic prompter to make generated co-speech gestures express semantics as well as align with speech.

A Simple Episodic Linear Probe Improves Visual Recognition in the Wild

Yuanzhi Liang, Linchao Zhu, et al. · IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022) · 2022 · pp. 9549-9559

ELP brings linear probing into training: an episodically reinitialized classifier learns on detached features to measure current discriminability, and ELP-SR uses the discrepancy between that probe and the main classifier to adaptively regularize samples.

2021 2

Food and Ingredient Joint Learning for Fine-Grained Recognition

Chengxu Liu, Yuanzhi Liang, et al. · IEEE Transactions on Circuits and Systems for Video Technology · 2021 · vol. 31(6) · pp. 2480-2493

The paper jointly learns food categories and ingredients with an Attention Fusion Network that emphasizes discriminative regions and a Food-Ingredient Joint Learning module using balance focal loss to address ingredient imbalance.

Removing Raindrops and Rain Streaks in One Go

Ruijie Quan, Xin Yu, et al. · IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2021) · 2021 · pp. 9143-9152

The method uses a complementary cascaded network with both raindrop→streak and streak→raindrop branches, fuses their outputs with attention, searches deraining blocks with neural architecture search, and introduces the real-world RainDS dataset.

2019 1