论文概要
VrR-VG 剪除仅凭非视觉统计就能预测的关系,从 Visual Genome 构建更强调视觉证据的 scene-graph 数据集,并联合编码实例、属性与关系来学习表示。
研究问题
视觉关系模型可能利用类别与频率偏差,在几乎不看图像证据的情况下预测 predicate,使 benchmark 分数无法真实反映视觉推理。
论文贡献
- 自动识别并移除 Visual Genome 中 visually irrelevant 的关系。
- 构建 VrR-VG,使统计捷径更难奏效、视觉证据更加必要。
- 联合实例、属性和关系学习 relationship-aware representation,并用于下游任务。
证据与评测范围
论文分析 VrR-VG 上 learnable method 与 statistical method 的差异,并报告学习特征对 image captioning 和 visual question answering 的系统性改善;准确幅度应引用 ICCV 表格。
适用范围与局限
剪除过程依赖论文对视觉无关性的定义和检测方式,一些关系本来就会结合视觉与上下文知识;数据集去偏也不能消除所有潜在 shortcut。
Related work 定位
VrR-VG 同时贡献数据集重构与关系表示学习。其通用去偏思路是:先检测哪些答案不用像素也能猜到,再构造更需要视觉证据的评测数据。
Related Work 表述
Liang 等从 Visual Genome 中剪除 visually irrelevant relationships 构建 VrR-VG,并联合建模实例、属性和关系来学习 relationship-aware feature。
这是一段用于说明论文定位的简洁中性表述。
论文官方英文摘要
Relationships encode the interactions among individual instances and play a critical role in deep visual scene understanding. Suffering from the high predictability with non-visual information, relationship models tend to fit the statistical bias rather than "learning" to infer the relationships from images. To encourage further development in visual relationships, we propose a novel method to mine more valuable relationships by automatically pruning visually-irrelevant relationships. We construct a new scene graph dataset named Visually-Relevant Relationships Dataset (VrR-VG) based on Visual Genome. Compared with existing datasets, the performance gap between learnable and statistical method is more significant in VrR-VG, and frequency-based analysis does not work anymore. Moreover, we propose to learn a relationship-aware representation by jointly considering instances, attributes and relationships. By applying the representation-aware feature learned on VrR-VG, the performances of image captioning and visual question answering are systematically improved, which demonstrates the effectiveness of both our dataset and features embedding schema. Both our VrR-VG dataset and representation-aware features will be made publicly available soon.
摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。
依据与出处
| 核对内容 | 论文中的位置 |
|---|---|
| 问题陈述 | Abstract |
| 方法与贡献 | Abstract; visually-irrelevant relationship pruning and relationship-aware representation sections |
| 评测结论 | Abstract; scene-graph analysis, image captioning, and VQA experiments |
主要核验来源: IEEE version of record (ICCV 2019 version of record; CVF open-access copy has alternate pagination 10403-10412; arXiv:1902.00313v2 cross-checked).
如何引用
科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。
Yuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian, Li Zhu, and Tao Mei. “VrR-VG: Refocusing Visually-Relevant Relationships.” 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (2019), 10402-10411. https://doi.org/10.1109/ICCV.2019.01050.
复用许可
本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.