SEEG: Semantic Energized Co-speech Gesture Generation
SEEG separates rhythmic and semantic cues in co-speech gesture, establishing that synchronized motion is not automatically communicative motion.
Research path
Studying how motion generation can satisfy semantic, relational, and physical constraints across human–human, human–object, and human–scene interaction.
Core question
This path studies three requirements for interaction-aware motion: semantic alignment between speech and gesture, action-conditioned affordance between a body and an object, and coordination among people, objects, and scenes. It then asks how generated interactions should be evaluated and connected to an executable embodied system.
A co-speech gesture can follow the rhythm while failing to express the semantic content of speech. A visible object part can be recognized even when the proposed contact and action are infeasible. Two motions can be plausible in isolation but inconsistent when performed together. These are different failures, but each shows that motion quality depends on context rather than kinematics alone.
The papers therefore address different conditional constraints on motion. SEEG uses speech semantics; MAAL conditions affordance on object state, contact, and robot action; InterSyn and Uni-Inter model coordination across multiple entities; LaxMotion studies the supervision needed for generalizable 3D motion. AntEval and the Embodied Brains roadmap extend the discussion from generation to process evaluation, execution, and verification.
SEEG separates beat-related and semantic information in co-speech gesture generation. Its Decoupled Mining module learns the two sources separately, while the Semantic Energizing Module adds supervision for semantic expression. The method therefore evaluates semantic content as a distinct objective rather than assuming that synchronization with speech rhythm is sufficient.
SEEG separates rhythmic and semantic cues in co-speech gesture, establishing that synchronized motion is not automatically communicative motion.
MAAL studies articulated-object affordance as a relation among object geometry, articulation state, contact location, and a candidate robot action. A visible object part is not assigned a fixed action label; the model estimates a multimodal distribution over interactions that are compatible with the current object and action conditions.
MAAL models articulated-object affordance as compatibility among geometry, object state, contact, and candidate robot action rather than appearance alone.
InterSyn jointly learns single-person and multi-person motion and refines relative coordination between participants. Uni-Inter encodes people, objects, and scenes in a Unified Interactive Volume and predicts motion within this shared spatial representation. The methods address different scopes, but both model an interaction through relations among entities rather than independent motion sequences.
InterSyn learns solo and multi-person dynamics together, then refines relative coordination so individual motion and mutual timing are not treated as separate problems.
Uni-Inter maps humans, objects, and scenes into a shared semantic occupancy volume, enabling one synthesis framework to reuse spatial knowledge across three interaction settings.
LaxMotion removes direct 3D pose regression and learns from global trajectories, monocular 2D kinematic cues, and structural regularization. The paper tests whether exact coordinate targets encourage fitting fixed training patterns and whether relaxed supervision can improve motion generation under distribution shift.
LaxMotion removes direct 3D pose regression and learns from trajectories, monocular cues, and structural regularization, questioning whether exact labels always improve generalization.
AntEval separates task completion from information exchange and intention expression in language-mediated multi-agent interaction. The Embodied Brains roadmap considers the corresponding systems problem in the physical world: model outputs must be translated into tool or controller requests, execution must be verified against the resulting state, and verified trajectories need explicit interfaces before they can be reused for learning.
AntEval evaluates task completion separately from information exchange and intention expression, showing that successful outcomes do not establish that agents communicated effectively.
The Embodied Brains roadmap connects predictive models to physical harnesses, shared contracts, verification, and experience reuse, defining a system horizon for grounded interaction.
These papers share a concern with context-dependent validity: whether motion expresses the intended semantics, is compatible with object state and action, coordinates across entities, or succeeds after execution. They do not share a common output space or evaluation protocol. SEEG generates co-speech gesture, MAAL estimates affordance, Uni-Inter generates 3D interaction motion, and AntEval evaluates language-mediated agents. Their grouping defines a research agenda, not a common model family.
Two gaps remain. Representation models must relate semantic intent and spatial context to feasible actions while expressing uncertainty about contact and dynamics. Evaluation must also move beyond offline motion quality to test whether an interaction can be executed, whether failure can be attributed to perception, prediction, planning, or control, and whether verified outcomes transfer across bodies, tasks, and environments.
This page connects collaborative research contributions; it does not replace the individual papers. Use each canonical paper record for evidence, source links, and citation downloads.