Summary
Uni-Inter represents human–human, human–object, and human–scene interactions in one Unified Interactive Volume and predicts motion probabilistically joint by joint, enabling one task-agnostic model to reason over heterogeneous and compound interaction contexts.
Research paths
How this paper contributes to the site's broader research map.
- Semantic Motion and Embodied InteractionCore work
Uni-Inter maps humans, objects, and scenes into a shared semantic occupancy volume, enabling one synthesis framework to reuse spatial knowledge across three interaction settings.
Research question
Interaction-motion systems are commonly designed per task, which fragments representations and limits transfer across people, objects, scenes, and combinations of these entities.
What the paper contributes
- Introduces a single architecture for human–human, human–object, and human–scene motion generation.
- Encodes heterogeneous entities in a shared volumetric field called the Unified Interactive Volume.
- Formulates generation as joint-wise probabilistic prediction to capture spatial dependencies and context-aware behavior.
Evidence and evaluation scope
Experiments cover three representative interaction tasks and report competitive performance plus generalization to novel entity combinations. Exact datasets, metrics, and comparisons should be taken from the ACM paper.
Scope and limitations
A shared representation does not imply that every interaction type or unseen composition is solved; demonstrated generalization is bounded by the evaluated tasks, entity encodings, and motion distributions.
Positioning for related work
Uni-Inter is a unified-representation approach to interaction-aware motion synthesis. Its contribution is not merely multi-task training: UIV supplies a common spatial field for relational reasoning across heterogeneous interaction types.
Related-work context
Liu et al. propose Uni-Inter, a task-agnostic 3D motion-generation framework that encodes human, object, and scene entities in a Unified Interactive Volume and performs joint-wise probabilistic motion prediction.
A concise, neutral description of how this paper can be situated in related work.
Official abstract
We present Uni-Inter, a unified framework for human motion generation that supports a wide range of interaction scenarios: including human-human, human-object, and human-scene-within a single, task-agnostic architecture. In contrast to existing methods that rely on task-specific designs and exhibit limited generalization, Uni-Inter introduces the Unified Interactive Volume (UIV), a volumetric representation that encodes heterogeneous interactive entities into a shared spatial field. This enables consistent relational reasoning and compound interaction modeling. Motion generation is formulated as joint-wise probabilistic prediction over the UIV, allowing the model to capture fine-grained spatial dependencies and produce coherent, context-aware behaviors. Experiments across three representative interaction tasks demonstrate that Uni-Inter achieves competitive performance and generalizes well to novel combinations of entities. These results suggest that unified modeling of compound interactions offers a promising direction for scalable motion synthesis in complex environments.
The abstract is reproduced for scholarly identification and remains under the paper publisher/authors’ original copyright; it is not covered by this page’s CC BY license.
Evidence references
| What to verify | Location in the paper |
|---|---|
| Problem statement | Abstract |
| Method and contributions | Abstract; Unified Interactive Volume and probabilistic prediction sections |
| Evaluation statement | Abstract; experiments across three interaction tasks |
Primary source: ACM version of record (SIGGRAPH Asia 2025 version of record; arXiv:2511.13032v1 used for accessible abstract).
How to cite
Cite the paper—not this explainer—for scientific claims. Cite this page only when reusing its original commentary.
Sheng Liu, Yuanzhi Liang, Jiepeng Wang, Sidan Du, Chi Zhang, and Xuelong Li. “Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts.” Proceedings of the SIGGRAPH Asia 2025 Conference Papers (2025), 1-11. https://doi.org/10.1145/3757377.3763954.
Reuse policy
Original explanatory text on this page is licensed under CC BY 4.0 with attribution and a link to this page. Paper title, abstract, figures, and bibliographic metadata are excluded and retain their original rights. Creative Commons Attribution 4.0 International.