Summary
VrR-VG prunes relationships that can be predicted from non-visual statistics, builds a visually relevant scene-graph dataset from Visual Genome, and learns representations that jointly encode instances, attributes, and relationships.
Research paths
How this paper contributes to the site's broader research map.
- Trustworthy Visual Generation Post-TrainingFoundation
VrR-VG uses a no-image relationship predictor to expose benchmark shortcuts, establishing the principle that performance must be tested against the information a model actually used.
Research question
Visual-relationship models can exploit category and frequency biases to predict predicates without using image evidence, making benchmark performance a weak test of visual reasoning.
What the paper contributes
- Automatically identifies and removes visually irrelevant relationships from Visual Genome.
- Constructs the VrR-VG dataset so statistical shortcuts are less effective and visual evidence matters more.
- Learns a relationship-aware representation over instances, attributes, and relations for downstream tasks.
Evidence and evaluation scope
The paper analyzes the gap between learnable and statistical methods on VrR-VG and reports systematic improvements in image captioning and visual question answering using the learned features. Exact margins belong to the ICCV tables.
Scope and limitations
Pruning is tied to the paper's definition and detector of visual irrelevance; some relationships can legitimately combine visual and contextual knowledge. Dataset debiasing does not eliminate every possible shortcut.
Positioning for related work
VrR-VG is both a dataset-refocusing and representation-learning contribution. It operationalizes a useful debiasing idea: first detect what can be guessed without pixels, then construct an evaluation set where image evidence is more necessary.
Related-work context
Liang et al. construct VrR-VG by pruning visually irrelevant relationships from Visual Genome and learn relationship-aware features that jointly model instances, attributes, and relations.
A concise, neutral description of how this paper can be situated in related work.
Official abstract
Relationships encode the interactions among individual instances and play a critical role in deep visual scene understanding. Suffering from the high predictability with non-visual information, relationship models tend to fit the statistical bias rather than "learning" to infer the relationships from images. To encourage further development in visual relationships, we propose a novel method to mine more valuable relationships by automatically pruning visually-irrelevant relationships. We construct a new scene graph dataset named Visually-Relevant Relationships Dataset (VrR-VG) based on Visual Genome. Compared with existing datasets, the performance gap between learnable and statistical method is more significant in VrR-VG, and frequency-based analysis does not work anymore. Moreover, we propose to learn a relationship-aware representation by jointly considering instances, attributes and relationships. By applying the representation-aware feature learned on VrR-VG, the performances of image captioning and visual question answering are systematically improved, which demonstrates the effectiveness of both our dataset and features embedding schema. Both our VrR-VG dataset and representation-aware features will be made publicly available soon.
The abstract is reproduced for scholarly identification and remains under the paper publisher/authors’ original copyright; it is not covered by this page’s CC BY license.
Evidence references
| What to verify | Location in the paper |
|---|---|
| Problem statement | Abstract |
| Method and contributions | Abstract; visually-irrelevant relationship pruning and relationship-aware representation sections |
| Evaluation statement | Abstract; scene-graph analysis, image captioning, and VQA experiments |
Primary source: IEEE version of record (ICCV 2019 version of record; CVF open-access copy has alternate pagination 10403-10412; arXiv:1902.00313v2 cross-checked).
How to cite
Cite the paper—not this explainer—for scientific claims. Cite this page only when reusing its original commentary.
Yuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian, Li Zhu, and Tao Mei. “VrR-VG: Refocusing Visually-Relevant Relationships.” 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (2019), 10402-10411. https://doi.org/10.1109/ICCV.2019.01050.
Reuse policy
Original explanatory text on this page is licensed under CC BY 4.0 with attribution and a link to this page. Paper title, abstract, figures, and bibliographic metadata are excluded and retain their original rights. Creative Commons Attribution 4.0 International.