论文概要
BPGO 不把所有视觉 reward score 视为同样可信,而是用 semantic prior 同时锚定组间 trust allocation 与组内 renormalization,使 GRPO 强调置信反馈并抑制歧义信号。
研究问题
文本和视觉内容之间是多对多关系,reward model 容易产生不确定或区分度弱的分数,标准 GRPO 可能因此低估可靠反馈并过拟合噪声比较。
论文贡献
- 引入 semantic prior anchor,显式建模视觉生成 reward 的不确定性。
- 根据各组与先验的一致性在 group 之间分配优化信任。
- 在 group 内重新归一化分数,放大可信偏差并压缩不确定分数。
证据与评测范围
论文报告在图像和视频生成中,相比标准 GRPO 及近期变体取得更强的语义对齐、感知质量和收敛速度。准确 baseline、指标和数值差异应引用正式实验。
适用范围与局限
BPGO 的行为依赖 semantic prior 是否合适以及底层 reward model 的质量。先验能够减少歧义,但不能证明所有保留信号在所有领域都符合人类判断。
Related work 定位
BPGO 是 GRPO 的 reward uncertainty 与 trust allocation 扩展。它区别于只提高 reward 粒度或只做时间 credit 的方法,重点校准哪些 group 和 sample 级反馈值得信任。
Related Work 表述
Liu 等提出 Bayesian Prior-Guided Optimization(BPGO),利用 semantic prior 在 GRPO group 之间分配信任并重整组内奖励,以提高图像和视频生成后训练信号的可靠性。
这是一段用于说明论文定位的简洁中性表述。
论文官方英文摘要
Group Relative Policy Optimization (GRPO) has emerged as an effective and lightweight framework for post-training visual generative models. However, its performance is fundamentally limited by the ambiguity of textual visual correspondence: a single prompt may validly describe diverse visual outputs, and a single image or video may support multiple equally correct interpretations. This many to many relationship leads reward models to generate uncertain and weakly discriminative signals, causing GRPO to underutilize reliable feedback and overfit noisy ones. We introduce Bayesian Prior-Guided Optimization (BPGO), a novel extension of GRPO that explicitly models reward uncertainty through a semantic prior anchor. BPGO adaptively modulates optimization trust at two levels: inter-group Bayesian trust allocation emphasizes updates from groups consistent with the prior while down-weighting ambiguous ones, and intra-group prior-anchored renormalization sharpens sample distinctions by expanding confident deviations and compressing uncertain scores. Across both image and video generation tasks, BPGO delivers consistently stronger semantic alignment, enhanced perceptual fidelity, and faster convergence than standard GRPO and recent variants.
摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。
依据与出处
| 核对内容 | 论文中的位置 |
|---|---|
| 问题陈述 | Abstract |
| 方法与贡献 | Abstract; inter-group trust allocation and intra-group renormalization sections |
| 评测结论 | Abstract; image and video experiments |
主要核验来源: CVF open-access paper (CVPR 2026 open-access version; arXiv:2511.18919v1 checked for first-posted metadata).
如何引用
科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。
Ruiying Liu, Yuanzhi Liang, Haibin Huang, Tianshu Yu, and Chi Zhang. “Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2026), 34408-34417.
复用许可
本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.