论文解读 · 作者已确认

VrR-VG: Refocusing Visually-Relevant Relationships

作者: , Yalong Bai, Wei Zhang, Xueming Qian, Li Zhu, Tao Mei

IEEE/CVF International Conference on Computer Vision (ICCV 2019) · 2019 · 页 10402-10411

visual relationshipsscene graphsdataset biasVisual Genomerepresentation learning

论文信息

作者
Yuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian, Li Zhu, and Tao Mei
推荐论文引用
Yuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian, Li Zhu, and Tao Mei. “VrR-VG: Refocusing Visually-Relevant Relationships.” 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (2019), 10402-10411. https://doi.org/10.1109/ICCV.2019.01050.
其他版本页码
CVF open-access copy: 10403-10412

版本日期

arXiv 首次提交
2019-02-01
arXiv 最近修订
2019-08-26
来源核验日期
2026-07-31
核验状态
作者已于 2026-07-31 核验

论文概要

VrR-VG 剪除仅凭非视觉统计就能预测的关系,从 Visual Genome 构建更强调视觉证据的 scene-graph 数据集,并联合编码实例、属性与关系来学习表示。

研究问题

视觉关系模型可能利用类别与频率偏差,在几乎不看图像证据的情况下预测 predicate,使 benchmark 分数无法真实反映视觉推理。

论文贡献

  • 自动识别并移除 Visual Genome 中 visually irrelevant 的关系。
  • 构建 VrR-VG,使统计捷径更难奏效、视觉证据更加必要。
  • 联合实例、属性和关系学习 relationship-aware representation,并用于下游任务。

证据与评测范围

论文分析 VrR-VG 上 learnable method 与 statistical method 的差异,并报告学习特征对 image captioning 和 visual question answering 的系统性改善;准确幅度应引用 ICCV 表格。

适用范围与局限

剪除过程依赖论文对视觉无关性的定义和检测方式,一些关系本来就会结合视觉与上下文知识;数据集去偏也不能消除所有潜在 shortcut。

Related work 定位

VrR-VG 同时贡献数据集重构与关系表示学习。其通用去偏思路是:先检测哪些答案不用像素也能猜到,再构造更需要视觉证据的评测数据。

Related Work 表述

Liang 等从 Visual Genome 中剪除 visually irrelevant relationships 构建 VrR-VG,并联合建模实例、属性和关系来学习 relationship-aware feature。

这是一段用于说明论文定位的简洁中性表述。

论文官方英文摘要

Relationships encode the interactions among individual instances and play a critical role in deep visual scene understanding. Suffering from the high predictability with non-visual information, relationship models tend to fit the statistical bias rather than "learning" to infer the relationships from images. To encourage further development in visual relationships, we propose a novel method to mine more valuable relationships by automatically pruning visually-irrelevant relationships. We construct a new scene graph dataset named Visually-Relevant Relationships Dataset (VrR-VG) based on Visual Genome. Compared with existing datasets, the performance gap between learnable and statistical method is more significant in VrR-VG, and frequency-based analysis does not work anymore. Moreover, we propose to learn a relationship-aware representation by jointly considering instances, attributes and relationships. By applying the representation-aware feature learned on VrR-VG, the performances of image captioning and visual question answering are systematically improved, which demonstrates the effectiveness of both our dataset and features embedding schema. Both our VrR-VG dataset and representation-aware features will be made publicly available soon.

摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。

依据与出处

核对内容论文中的位置
问题陈述Abstract
方法与贡献Abstract; visually-irrelevant relationship pruning and relationship-aware representation sections
评测结论Abstract; scene-graph analysis, image captioning, and VQA experiments

主要核验来源: IEEE version of record (ICCV 2019 version of record; CVF open-access copy has alternate pagination 10403-10412; arXiv:1902.00313v2 cross-checked).

如何引用

科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。

Yuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian, Li Zhu, and Tao Mei. “VrR-VG: Refocusing Visually-Relevant Relationships.” 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (2019), 10402-10411. https://doi.org/10.1109/ICCV.2019.01050.

复用许可

本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.

一手来源与资源