Paper overview · author-verified

SEEG: Semantic Energized Co-speech Gesture Generation

Authors: , Qianyu Feng, Linchao Zhu, Li Hu, Pan Pan, Yi Yang

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022) · 2022 · pp. 10463-10472

co-speech gesturegesture generationsemantic gesturesspeech rhythmdisentangled learning

Publication details

Authors
Yuanzhi Liang, Qianyu Feng, Linchao Zhu, Li Hu, Pan Pan, and Yi Yang
Recommended paper citation
Yuanzhi Liang, Qianyu Feng, Linchao Zhu, Li Hu, Pan Pan, and Yi Yang. “SEEG: Semantic Energized Co-speech Gesture Generation.” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 10463-10472. https://doi.org/10.1109/CVPR52688.2022.01022.
Alternate-copy pagination
CVF open-access copy: 10473-10482

Version dates

Source checked
2026-07-31
Verification status
Author-verified on 2026-07-31

Summary

SEEG separates beat-related and semantic information with a DEcoupled Mining module, then uses a Semantic Energizing Module and semantic prompter to make generated co-speech gestures express semantics as well as align with speech.

Research paths

How this paper contributes to the site's broader research map.

Research question

Co-speech gesture models can learn rhythmic alignment while failing to explicitly capture and express the semantic content conveyed by meaningful gestures.

What the paper contributes

  • Introduces DEcoupled Mining (DEM) to separate information for beat gestures and semantic gestures.
  • Introduces a Semantic Energizing Module (SEM) that supervises semantic expression, not only representation similarity.
  • Uses a semantic prompter to transfer semantic-aware supervision to generated gestures.

Evidence and evaluation scope

The paper reports results on multiple benchmarks and three metrics, with gains on semantic-aware evaluations plus qualitative improvements in expressiveness. Exact metrics and dataset results should be cited from the CVPR paper.

Scope and limitations

The semantic categories and supervision available to SEM shape what counts as expressive meaning. The method does not imply that all culturally dependent gesture semantics or open-domain communicative intent are captured.

Positioning for related work

SEEG is a semantic-aware co-speech gesture method that explicitly separates easier rhythmic cues from harder semantic cues and adds supervision for semantic expression.

Related-work context

Liang et al. propose SEEG, which decouples beat and semantic gesture cues and applies semantic-aware supervision through a Semantic Energizing Module for co-speech gesture generation.

A concise, neutral description of how this paper can be situated in related work.

Official abstract

Talking gesture generation is a practical yet challenging task which aims to synthesize gestures in line with speech. Gestures with meaningful signs can better convey useful information and arouse sympathy in the audience. Current works focus on aligning gestures with the speech rhythms, which are hard to mine the semantics and model semantic gestures explicitly. In this paper, we propose a novel method SEmantic Energized Generation (SEEG), for semantic-aware gesture generation. Our method contains two parts: DEcoupled Mining module (DEM) and Semantic Energizing Module (SEM). DEM decouples the semantic-irrelevant information from inputs and separately mines information for the beat and semantic gestures. SEM conducts semantic learning and produces semantic gestures. Apart from representational similarity, SEM requires the predictions to express the same semantics as the ground truth. Besides, a semantic prompter is designed in SEM to leverage the semantic-aware supervision to predictions. This promotes the networks to learn and generate semantic gestures. Experimental results reported in three metrics on different benchmarks prove that SEEG efficiently mines semantic cues and generates semantic gestures. In comparison, SEEG outperforms other methods in all semantic-aware evaluations on different datasets. Qualitative evaluations also indicate the superiority of SEEG in semantic expressiveness.

The abstract is reproduced for scholarly identification and remains under the paper publisher/authors’ original copyright; it is not covered by this page’s CC BY license.

Evidence references

What to verifyLocation in the paper
Problem statementAbstract
Method and contributionsAbstract; DEcoupled Mining and Semantic Energizing Module sections
Evaluation statementAbstract; semantic-aware quantitative and qualitative evaluations

Primary source: IEEE version of record (CVPR 2022 version of record; CVF open-access copy has alternate pagination 10473-10482).

How to cite

Cite the paper—not this explainer—for scientific claims. Cite this page only when reusing its original commentary.

Yuanzhi Liang, Qianyu Feng, Linchao Zhu, Li Hu, Pan Pan, and Yi Yang. “SEEG: Semantic Energized Co-speech Gesture Generation.” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 10463-10472. https://doi.org/10.1109/CVPR52688.2022.01022.

Reuse policy

Original explanatory text on this page is licensed under CC BY 4.0 with attribution and a link to this page. Paper title, abstract, figures, and bibliographic metadata are excluded and retain their original rights. Creative Commons Attribution 4.0 International.

Primary sources and resources