论文解读 · 作者已确认

Rethinking Reward Signals in Video GRPO: When Scores Become Targets

作者: Rui Li, , Ziqi Ni, Haibin Huang, Chi Zhang, Xuelong Li

European Conference on Computer Vision (ECCV 2026, accepted) · 2026 · 待正式出版

video generationGRPOreward saturationreward hackingGoodhart's law

论文信息

作者
Rui Li, Yuanzhi Liang, Ziqi Ni, Haibin Huang, Chi Zhang, and Xuelong Li
推荐论文引用
Rui Li, Yuanzhi Liang, Ziqi Ni, Haibin Huang, Chi Zhang, and Xuelong Li. “Rethinking Reward Signals in Video GRPO: When Scores Become Targets.” European Conference on Computer Vision (ECCV 2026) (2026, forthcoming). arXiv:2511.19356.
书目说明
The status “accepted at ECCV 2026” is author-supplied. The paper-level Springer/ECVA proceedings record, DOI, volume, and pagination were not yet public on 2026-07-31; this export is explicitly marked forthcoming and includes the current arXiv identifier.

版本日期

arXiv 首次提交
2025-11-24
arXiv 最近修订
2026-07-17
来源核验日期
2026-07-31
核验状态
作者已于 2026-07-31 核验

论文概要

TaRoS v4 针对 video GRPO 中的 reward hacking 与饱和:用 component-level performance assessment 和 intra-group sparsity 组织多维反馈,并自适应降低已经饱和的 reward component 权重。

研究问题

GRPO 持续优化 reward-induced advantage 后,reward score 可能不再忠实代表真实视频质量:复合目标会诱发 shortcut,prompt group 内的分数也可能饱和并失去排序信息。

论文贡献

  • 在 reward component 层面评估表现,而不是把复合分数视为始终稳定的目标。
  • 利用 intra-group sparsity 将多维奖励组织到优化目标上。
  • 自适应降低饱和 component 的权重,保留有效优化方向和组内排序差异。

证据与评测范围

当前 arXiv v4 报告相对强 baseline 在视觉质量、运动一致性和文本—视频对齐上的一致改进。针对旧版本的笼统“目标校准”描述不应继续用于 v4。

适用范围与局限

TaRoS 只能通过已有 reward component 诊断问题,未被奖励覆盖或系统性偏置的维度仍无法自动修正;当前公开记录是与 ECCV 2026 accepted paper 关联的修订预印本,最终 proceedings 信息可能变化。

Related work 定位

TaRoS 属于视频生成 RL 的 robust reward signaling 工作。它不只是聚合奖励,而是直接处理持续 GRPO 优化中的 Goodhart-style shortcut 与 component saturation。

Related Work 表述

Li 等提出 TaRoS,一种面向 video GRPO 的 target-robust reward-signaling 框架,通过 component-level assessment、intra-group sparsity 和对饱和奖励分量的自适应降权减少 reward hacking。

这是一段用于说明论文定位的简洁中性表述。

论文官方英文摘要

Group Relative Policy Optimization (GRPO) enables stable and preference-oriented updates via group-wise comparisons for post-training video generation. However, GRPO directly optimizes reward-induced advantages. Under sustained optimization, the reward score can lose fidelity as a proxy for true video quality, consistent with the phenomenon described by Goodhart's Law. This leads to two recurring issues: (i) shortcut-driven optimization under composite objectives and (ii) reward saturation within prompt groups. To address these issues, we introduce TaRoS, a Target-Robust Reward Signaling framework for Video generation GRPO. TaRoS leverages component level performance assessment together with intra-group sparsity to organize multi-aspect rewards towards optimization objectives. In addition, it adaptively downweights components that exhibit saturation, thereby preserving effective optimization directions and mitigating redundancy. This maintains meaningful optimization directions and preserves within-group ranking separation, thereby preventing reward hacking and leading to more reliable policy updates. Extensive experiments show consistent improvements in visual fidelity, motion coherence, and text-video alignment over strong baselines.

摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。

依据与出处

核对内容论文中的位置
问题陈述Abstract
方法与贡献Abstract; TaRoS sections on component-level assessment, intra-group sparsity, and saturation downweighting
评测结论Abstract; video generation experiments in v4

主要核验来源: arXiv record (arXiv:2511.19356v4, revised 2026-07-17).

如何引用

科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。

Rui Li, Yuanzhi Liang, Ziqi Ni, Haibin Huang, Chi Zhang, and Xuelong Li. “Rethinking Reward Signals in Video GRPO: When Scores Become Targets.” European Conference on Computer Vision (ECCV 2026) (2026, forthcoming). arXiv:2511.19356.

复用许可

本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.

一手来源与资源