论文概要
OTCA 不再把一个标量 reward 平均传给所有去噪步,而是同时回答两个问题:轨迹中的哪些 step 更重要,以及每个阶段应该强调哪些 reward objective。
研究问题
视觉 GRPO 往往把画质、运动一致性、文本对齐等异构奖励压成一个静态标量,并对整个 diffusion trajectory 均匀传播,忽略了不同去噪阶段的职责差异。
论文贡献
- Trajectory-Level Credit Decomposition 估计不同去噪 step 的相对重要性。
- Multi-Objective Credit Allocation 在去噪过程中自适应组合多个奖励目标。
- 两者联合把粗粒度 reward 转化为同时具备 timestep awareness 与 objective awareness 的训练信号。
证据与评测范围
论文报告了在图像和视频生成任务及多项评测指标上的一致改进。本页不转述具体数值,模型、数据集与指标层面的比较应以论文实验表格为准。
适用范围与局限
方法面向 diffusion visual generator 的 GRPO 后训练,仍然依赖底层 reward model 的覆盖范围和可靠性;会议正式 proceedings 发布后还需更新最终书目信息。
Related work 定位
OTCA 适合被定位为视觉生成 RL 的细粒度 credit assignment 方法。它不是只做多奖励聚合或样本重权重,而是同时建模时间维度与目标维度的 credit。
Related Work 表述
Li 等提出 Objective-aware Trajectory Credit Assignment(OTCA),通过分解不同去噪步的贡献并自适应分配多种奖励目标,为基于 GRPO 的图像和视频生成提供结构化训练信号。
这是一段用于说明论文定位的简洁中性表述。
论文官方英文摘要
Reinforcement learning, particularly Group Relative Policy Optimization (GRPO), has emerged as an effective framework for post-training visual generative models with human preference signals. However, its effectiveness is fundamentally limited by coarse reward credit assignment. In modern visual generation, multiple reward models are often used to capture heterogeneous objectives, such as visual quality, motion consistency, and text alignment. Existing GRPO pipelines typically collapse these rewards into a single static scalar and propagate it uniformly across the entire diffusion trajectory. This design ignores the stage-specific roles of different denoising steps and produces mistimed or incompatible optimization signals. To address this issue, we propose Objective-aware Trajectory Credit Assignment (OTCA), a structured framework for fine-grained GRPO training. OTCA consists of two key components. Trajectory-Level Credit Decomposition estimates the relative importance of different denoising steps. Multi-Objective Credit Allocation adaptively weights and combines multiple reward signals throughout the denoising process. By jointly modeling temporal credit and objective-level credit, OTCA converts coarse reward supervision into a structured, timestep-aware training signal that better matches the iterative nature of diffusion-based generation. Extensive experiments show that OTCA consistently improves both image and video generation quality across evaluation metrics.
摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。
依据与出处
| 核对内容 | 论文中的位置 |
|---|---|
| 问题陈述 | Abstract |
| 方法与贡献 | Abstract; method sections on Trajectory-Level Credit Decomposition and Multi-Objective Credit Allocation |
| 评测结论 | Abstract; image and video experiments |
主要核验来源: arXiv record (arXiv:2604.19234v2).
如何引用
科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。
Rui Li, Ke Hao, Yuanzhi Liang, Haibin Huang, Chi Zhang, Yun Gu, and XueLong Li. “Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation.” 34th ACM International Conference on Multimedia (ACM Multimedia 2026) (2026, forthcoming). arXiv:2604.19234.
复用许可
本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.