论文概要
TaRoS v4 针对 video GRPO 中的 reward hacking 与饱和:用 component-level performance assessment 和 intra-group sparsity 组织多维反馈,并自适应降低已经饱和的 reward component 权重。
研究问题
GRPO 持续优化 reward-induced advantage 后,reward score 可能不再忠实代表真实视频质量:复合目标会诱发 shortcut,prompt group 内的分数也可能饱和并失去排序信息。
论文贡献
- 在 reward component 层面评估表现,而不是把复合分数视为始终稳定的目标。
- 利用 intra-group sparsity 将多维奖励组织到优化目标上。
- 自适应降低饱和 component 的权重,保留有效优化方向和组内排序差异。
证据与评测范围
当前 arXiv v4 报告相对强 baseline 在视觉质量、运动一致性和文本—视频对齐上的一致改进。针对旧版本的笼统“目标校准”描述不应继续用于 v4。
适用范围与局限
TaRoS 只能通过已有 reward component 诊断问题,未被奖励覆盖或系统性偏置的维度仍无法自动修正;当前公开记录是与 ECCV 2026 accepted paper 关联的修订预印本,最终 proceedings 信息可能变化。
Related work 定位
TaRoS 属于视频生成 RL 的 robust reward signaling 工作。它不只是聚合奖励,而是直接处理持续 GRPO 优化中的 Goodhart-style shortcut 与 component saturation。
Related Work 表述
Li 等提出 TaRoS,一种面向 video GRPO 的 target-robust reward-signaling 框架,通过 component-level assessment、intra-group sparsity 和对饱和奖励分量的自适应降权减少 reward hacking。
这是一段用于说明论文定位的简洁中性表述。
论文官方英文摘要
Group Relative Policy Optimization (GRPO) enables stable and preference-oriented updates via group-wise comparisons for post-training video generation. However, GRPO directly optimizes reward-induced advantages. Under sustained optimization, the reward score can lose fidelity as a proxy for true video quality, consistent with the phenomenon described by Goodhart's Law. This leads to two recurring issues: (i) shortcut-driven optimization under composite objectives and (ii) reward saturation within prompt groups. To address these issues, we introduce TaRoS, a Target-Robust Reward Signaling framework for Video generation GRPO. TaRoS leverages component level performance assessment together with intra-group sparsity to organize multi-aspect rewards towards optimization objectives. In addition, it adaptively downweights components that exhibit saturation, thereby preserving effective optimization directions and mitigating redundancy. This maintains meaningful optimization directions and preserves within-group ranking separation, thereby preventing reward hacking and leading to more reliable policy updates. Extensive experiments show consistent improvements in visual fidelity, motion coherence, and text-video alignment over strong baselines.
摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。
依据与出处
| 核对内容 | 论文中的位置 |
|---|---|
| 问题陈述 | Abstract |
| 方法与贡献 | Abstract; TaRoS sections on component-level assessment, intra-group sparsity, and saturation downweighting |
| 评测结论 | Abstract; video generation experiments in v4 |
主要核验来源: arXiv record (arXiv:2511.19356v4, revised 2026-07-17).
如何引用
科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。
Rui Li, Yuanzhi Liang, Ziqi Ni, Haibin Huang, Chi Zhang, and Xuelong Li. “Rethinking Reward Signals in Video GRPO: When Scores Become Targets.” European Conference on Computer Vision (ECCV 2026) (2026, forthcoming). arXiv:2511.19356.
复用许可
本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.