论文概要
ViPO 把每个生成样本的单一标量 reward 转换为空间和时间结构化的像素级 advantage map,利用预训练视觉特征把 GRPO 更新集中到感知上重要的区域,同时保留标准训练流程。
研究问题
标量奖励把整张图像或整段视频视为不可分割的整体,GRPO 因而难以针对局部伪影或细粒度时空线索进行定向优化。
论文贡献
- 提出把标量反馈提升为像素级结构化 advantage 的 GRPO 变体。
- 利用预训练视觉 backbone 构建具有空间和时间感知能力的 advantage map。
- 保持 architecture-agnostic、轻量并兼容现有 GRPO pipeline。
证据与评测范围
论文报告在图像和视频 benchmark 上优于 vanilla GRPO,同时改善域内偏好对齐与域外泛化。准确奖励、数据集、backbone 和提升幅度应以正式论文表格为准。
适用范围与局限
结构化 advantage map 由预训练视觉特征诱导,因此受这些 backbone 表征能力的影响;像素级优化也不能自动保证每个 reward model 都可靠或具备真实的局部因果性。
Related work 定位
ViPO 属于提高生成模型 RL 反馈粒度的工作。其关键不是替换生成器或改变 GRPO 总体框架,而是对 advantage 进行时空重分配。
Related Work 表述
Ni 等提出 Visual Preference Policy Optimization(ViPO),通过 perceptual structuring module 将标量奖励转换为具备时空感知的像素级优势,用于基于 GRPO 的图像和视频生成。
这是一段用于说明论文定位的简洁中性表述。
论文官方英文摘要
Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generators with human preferences. However, existing GRPO pipelines rely on a single scalar reward per sample, treating each image or video as a holistic entity and ignoring the rich spatial and temporal structure of visual content. This coarse supervision hinders the correction of localized artifacts and the modeling of fine-grained perceptual cues. We introduce Visual Preference Policy Optimization (ViPO), a GRPO variant that lifts scalar feedback into structured, pixel-level advantages. ViPO employs a Perceptual Structuring Module that uses pretrained vision backbones to construct spatially and temporally aware advantage maps, redistributing optimization pressure toward perceptually important regions while preserving the stability of standard GRPO. Across both image and video benchmarks, ViPO consistently outperforms vanilla GRPO, improving in-domain alignment with human-preference rewards and enhancing generalization on out-of-domain evaluations. The method is architecture-agnostic, lightweight, and fully compatible with existing GRPO training pipelines, providing a more expressive and informative learning signal for visual generation.
摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。
依据与出处
| 核对内容 | 论文中的位置 |
|---|---|
| 问题陈述 | Abstract |
| 方法与贡献 | Abstract; Perceptual Structuring Module section |
| 评测结论 | Abstract; image and video benchmark sections |
主要核验来源: CVF open-access paper (CVPR 2026 open-access version; arXiv:2511.18719v4 checked for revision date).
如何引用
科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。
Ziqi Ni, Yuanzhi Liang, Rui Li, Yi Zhou, Haibin Huang, Chi Zhang, and Xuelong Li. “Seeing What Matters: Visual Preference Policy Optimization for Visual Generation.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2026), 27260-27269.
复用许可
本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.