DeepAI AI Chat
Log In Sign Up

Rethinking Visual Relationships for High-level Image Understanding

by   Yuanzhi Liang, et al.

Relationships, as the bond of isolated entities in images, reflect the interaction between objects and lead to a semantic understanding of scenes. Suffering from visually-irrelevant relationships in current scene graph datasets, the utilization of relationships for semantic tasks is difficult. The datasets widely used in scene graph generation tasks are splitted from Visual Genome by label frequency, which even can be well solved by statistical counting. To encourage further development in relationships, we propose a novel method to mine more valuable relationships by automatically filtering out visually-irrelevant relationships. Then, we construct a new scene graph dataset named Visually-Relevant Relationships Dataset (VrR-VG) from Visual Genome. We evaluate several existing methods in scene graph generation in our dataset. The results show the performances degrade significantly compared to the previous dataset and the frequency analysis do not work on our dataset anymore. Moreover, we propose a method to learn feature representations of instances, attributes, and visual relationships jointly from images, then we apply the learned features to image captioning and visual question answering respectively. The improvements on the both tasks demonstrate the efficiency of the features with relation information and the richer semantic information provided in our dataset.


page 1

page 2

page 6

page 11

page 12


Scene Graph Generation for Better Image Captioning?

We investigate the incorporation of visual relationships into the task o...

Scene Graph Generation with Geometric Context

Scene Graph Generation has gained much attention in computer vision rese...

Semantic Compositional Learning for Low-shot Scene Graph Generation

Scene graphs provide valuable information to many downstream tasks. Many...

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Despite progress in perceptual tasks such as image classification, compu...

Attentive Relational Networks for Mapping Images to Scene Graphs

Scene graph generation refers to the task of automatically mapping an im...

Improving Information Extraction from Images with Learned Semantic Models

Many applications require an understanding of an image that goes beyond ...

Probabilistic Modeling of Semantic Ambiguity for Scene Graph Generation

To generate "accurate" scene graphs, almost all existing methods predict...