Paper overview · author-verified

Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Authors: Rui Li, , Ziqi Ni, Haibin Huang, Chi Zhang, Xuelong Li

European Conference on Computer Vision (ECCV 2026, accepted) · 2026 · forthcoming

video generationGRPOreward saturationreward hackingGoodhart's law

Publication details

Authors
Rui Li, Yuanzhi Liang, Ziqi Ni, Haibin Huang, Chi Zhang, and Xuelong Li
Recommended paper citation
Rui Li, Yuanzhi Liang, Ziqi Ni, Haibin Huang, Chi Zhang, and Xuelong Li. “Rethinking Reward Signals in Video GRPO: When Scores Become Targets.” European Conference on Computer Vision (ECCV 2026) (2026, forthcoming). arXiv:2511.19356.
Bibliographic note
The status “accepted at ECCV 2026” is author-supplied. The paper-level Springer/ECVA proceedings record, DOI, volume, and pagination were not yet public on 2026-07-31; this export is explicitly marked forthcoming and includes the current arXiv identifier.

Version dates

arXiv first posted
2025-11-24
arXiv last revised
2026-07-17
Source checked
2026-07-31
Verification status
Author-verified on 2026-07-31

Summary

TaRoS v4 addresses reward hacking and saturation in video GRPO by organizing multi-aspect feedback with component-level performance assessment and intra-group sparsity, then adaptively downweighting components whose scores have saturated.

Research paths

How this paper contributes to the site's broader research map.

Research question

When GRPO repeatedly optimizes reward-induced advantages, reward scores can stop being faithful proxies for true video quality: composite objectives invite shortcuts and scores within a prompt group may saturate, erasing useful rankings.

What the paper contributes

  • Assesses performance at the reward-component level rather than treating a composite score as one stable target.
  • Uses intra-group sparsity to organize multi-aspect rewards toward optimization objectives.
  • Adaptively downweights saturated components to preserve effective directions and within-group ranking separation.

Evidence and evaluation scope

The current arXiv v4 reports consistent improvements in visual fidelity, motion coherence, and text–video alignment over strong baselines. Earlier descriptions centered on generic target calibration should not be used for this version.

Scope and limitations

TaRoS diagnoses behavior through the available reward components, so missing or systematically biased aspects remain outside its correction mechanism. The public record is a revised preprint associated with an accepted ECCV 2026 paper; final proceedings metadata may change.

Positioning for related work

TaRoS belongs to robust reward signaling for video-generation RL. Unlike reward aggregation alone, it explicitly responds to Goodhart-style shortcut optimization and component saturation during continued GRPO training.

Related-work context

Li et al. propose TaRoS, a target-robust reward-signaling framework for video GRPO that combines component-level assessment, intra-group sparsity, and adaptive downweighting of saturated reward components to reduce reward hacking.

A concise, neutral description of how this paper can be situated in related work.

Official abstract

Group Relative Policy Optimization (GRPO) enables stable and preference-oriented updates via group-wise comparisons for post-training video generation. However, GRPO directly optimizes reward-induced advantages. Under sustained optimization, the reward score can lose fidelity as a proxy for true video quality, consistent with the phenomenon described by Goodhart's Law. This leads to two recurring issues: (i) shortcut-driven optimization under composite objectives and (ii) reward saturation within prompt groups. To address these issues, we introduce TaRoS, a Target-Robust Reward Signaling framework for Video generation GRPO. TaRoS leverages component level performance assessment together with intra-group sparsity to organize multi-aspect rewards towards optimization objectives. In addition, it adaptively downweights components that exhibit saturation, thereby preserving effective optimization directions and mitigating redundancy. This maintains meaningful optimization directions and preserves within-group ranking separation, thereby preventing reward hacking and leading to more reliable policy updates. Extensive experiments show consistent improvements in visual fidelity, motion coherence, and text-video alignment over strong baselines.

The abstract is reproduced for scholarly identification and remains under the paper publisher/authors’ original copyright; it is not covered by this page’s CC BY license.

Evidence references

What to verifyLocation in the paper
Problem statementAbstract
Method and contributionsAbstract; TaRoS sections on component-level assessment, intra-group sparsity, and saturation downweighting
Evaluation statementAbstract; video generation experiments in v4

Primary source: arXiv record (arXiv:2511.19356v4, revised 2026-07-17).

How to cite

Cite the paper—not this explainer—for scientific claims. Cite this page only when reusing its original commentary.

Rui Li, Yuanzhi Liang, Ziqi Ni, Haibin Huang, Chi Zhang, and Xuelong Li. “Rethinking Reward Signals in Video GRPO: When Scores Become Targets.” European Conference on Computer Vision (ECCV 2026) (2026, forthcoming). arXiv:2511.19356.

Reuse policy

Original explanatory text on this page is licensed under CC BY 4.0 with attribution and a link to this page. Paper title, abstract, figures, and bibliographic metadata are excluded and retain their original rights. Creative Commons Attribution 4.0 International.

Primary sources and resources