论文概要
RATS 把轨迹蒸馏与偏好反馈结合起来:horizon matching 在关键去噪阶段对齐 teacher 与 student,reward-aware gate 只在 teacher 的奖励表现更好时加强指导。
研究问题
固定蒸馏要求 few-step student 持续模仿 multi-step teacher,容易把 teacher 变成性能上限,即使偏好奖励表明 student 已经能够达到或超过 teacher。
论文贡献
- 通过 horizon matching 对齐 teacher 和 student 在关键阶段的 latent trajectory。
- 根据双方相对 reward 表现动态增强或减弱 teacher guidance。
- 在不增加推理开销的前提下结合轨迹蒸馏与 preference alignment。
证据与评测范围
论文报告 few-step 生成在效率—质量权衡上获得改进,并缩小与强 multi-step generator 的差距。具体步数、底模、奖励和指标数值应查阅当前版本的实验表格。
适用范围与局限
Reward-aware gate 会继承用于比较 teacher 与 student 的奖励信号所包含的假设与偏差;方法面向 few-step visual generation,迁移到其他蒸馏设置仍需验证。
Related work 定位
RATS 位于 trajectory distillation 与 reward-based alignment 的交叉点。它的关键差异是根据相对奖励质量决定是否跟随 teacher,而不是始终把 teacher 当作固定目标。
Related Work 表述
Li 等提出 Reward-Aware Trajectory Shaping(RATS),利用 horizon matching 对齐师生轨迹,并通过 reward-aware gate 有条件地调节 teacher guidance,以实现偏好对齐的少步视觉生成。
这是一段用于说明论文定位的简洁中性表述。
论文官方英文摘要
Achieving high-fidelity generation in extremely few sampling steps has long been a central goal of generative modeling. Existing approaches largely rely on distillation-based frameworks to compress the original multi-step denoising process into a few-step generator. However, such methods inherently constrain the student to imitate a stronger multi-step teacher, imposing the teacher as an upper bound on student performance. We argue that introducing preference alignment awareness enables the student to optimize toward reward-preferred generation quality, potentially surpassing the teacher instead of being restricted to rigid teacher imitation. To this end, we propose Reward-Aware Trajectory Shaping (RATS), a lightweight framework for preference-aligned few-step generation. Specifically, teacher and student latent trajectories are aligned at key denoising stages through horizon matching, while a reward-aware gate is introduced to adaptively regulate teacher guidance based on their relative reward performance. Trajectory shaping is strengthened when the teacher achieves higher rewards, and relaxed when the student matches or surpasses the teacher, thereby enabling continued reward-driven improvement. By seamlessly integrating trajectory distillation, reward-aware gating, and preference alignment, RATS effectively transfers preference-relevant knowledge from high-step generators without incurring additional test-time computational overhead. Experimental results demonstrate that RATS substantially improves the efficiency--quality trade-off in few-step visual generation, significantly narrowing the gap between few-step students and stronger multi-step generators.
摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。
依据与出处
| 核对内容 | 论文中的位置 |
|---|---|
| 问题陈述 | Abstract |
| 方法与贡献 | Abstract; method sections on horizon matching and the reward-aware gate |
| 评测结论 | Abstract; few-step generation experiments |
主要核验来源: arXiv record (arXiv:2604.14910v3).
如何引用
科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。
Rui Li, Bingyu Li, Yuanzhi Liang, Haibin Huang, Chi Zhang, and XueLong Li. “Reward-Aware Trajectory Shaping for Few-step Visual Generation.” 34th ACM International Conference on Multimedia (ACM Multimedia 2026) (2026, forthcoming). arXiv:2604.14910.
复用许可
本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.