Research path

Semantic Motion and Embodied Interaction

Studying how motion generation can satisfy semantic, relational, and physical constraints across human–human, human–object, and human–scene interaction.

中文

Core question

What representations and evaluation criteria are needed to generate motion that conveys the intended semantics, remains physically feasible, and coordinates with other entities?

This path studies three requirements for interaction-aware motion: semantic alignment between speech and gesture, action-conditioned affordance between a body and an object, and coordination among people, objects, and scenes. It then asks how generated interactions should be evaluated and connected to an executable embodied system.

A co-speech gesture can follow the rhythm while failing to express the semantic content of speech. A visible object part can be recognized even when the proposed contact and action are infeasible. Two motions can be plausible in isolation but inconsistent when performed together. These are different failures, but each shows that motion quality depends on context rather than kinematics alone.

The papers therefore address different conditional constraints on motion. SEEG uses speech semantics; MAAL conditions affordance on object state, contact, and robot action; InterSyn and Uni-Inter model coordination across multiple entities; LaxMotion studies the supervision needed for generalizable 3D motion. AntEval and the Embodied Brains roadmap extend the discussion from generation to process evaluation, execution, and verification.

Separate rhythmic alignment from semantic expression

SEEG separates beat-related and semantic information in co-speech gesture generation. Its Decoupled Mining module learns the two sources separately, while the Semantic Energizing Module adds supervision for semantic expression. The method therefore evaluates semantic content as a distinct objective rather than assuming that synchronization with speech rhythm is sufficient.

Condition affordance on object state and action

MAAL studies articulated-object affordance as a relation among object geometry, articulation state, contact location, and a candidate robot action. A visible object part is not assigned a fixed action label; the model estimates a multimodal distribution over interactions that are compatible with the current object and action conditions.

Represent coordination across interaction settings

InterSyn jointly learns single-person and multi-person motion and refines relative coordination between participants. Uni-Inter encodes people, objects, and scenes in a Unified Interactive Volume and predicts motion within this shared spatial representation. The methods address different scopes, but both model an interaction through relations among entities rather than independent motion sequences.

Study supervision for motion generalization

LaxMotion removes direct 3D pose regression and learns from global trajectories, monocular 2D kinematic cues, and structural regularization. The paper tests whether exact coordinate targets encourage fitting fixed training patterns and whether relaxed supervision can improve motion generation under distribution shift.

Evaluate interaction and verify execution

AntEval separates task completion from information exchange and intention expression in language-mediated multi-agent interaction. The Embodied Brains roadmap considers the corresponding systems problem in the physical world: model outputs must be translated into tool or controller requests, execution must be verified against the resulting state, and verified trajectories need explicit interfaces before they can be reused for learning.

What connects these papers—and what remains different

These papers share a concern with context-dependent validity: whether motion expresses the intended semantics, is compatible with object state and action, coordinates across entities, or succeeds after execution. They do not share a common output space or evaluation protocol. SEEG generates co-speech gesture, MAAL estimates affordance, Uni-Inter generates 3D interaction motion, and AntEval evaluates language-mediated agents. Their grouping defines a research agenda, not a common model family.

Open questions

Two gaps remain. Representation models must relate semantic intent and spatial context to feasible actions while expressing uncertainty about contact and dynamics. Evaluation must also move beyond offline motion quality to test whether an interaction can be executed, whether failure can be attributed to perception, prediction, planning, or control, and whether verified outcomes transfer across bodies, tasks, and environments.

Read and cite the papers

This page connects collaborative research contributions; it does not replace the individual papers. Use each canonical paper record for evidence, source links, and citation downloads.