Utilizing Large Scale Vision and Text Datasets for Image Segmentation from Referring Expressions

08/30/2016
by   Ronghang Hu, et al.
0

Image segmentation from referring expressions is a joint vision and language modeling task, where the input is an image and a textual expression describing a particular region in the image; and the goal is to localize and segment the specific image region based on the given expression. One major difficulty to train such language-based image segmentation systems is the lack of datasets with joint vision and text annotations. Although existing vision datasets such as MS COCO provide image captions, there are few datasets with region-level textual annotations for images, and these are often smaller in scale. In this paper, we explore how existing large scale vision-only and text-only datasets can be utilized to train models for image segmentation from referring expressions. We propose a method to address this problem, and show in experiments that our method can help this joint vision and language modeling task with vision-only and text-only data and outperforms previous results.

READ FULL TEXT

page 1

page 4

page 7

page 8

page 9

page 10

research
09/25/2019

UNITER: Learning UNiversal Image-TExt Representations

Joint image-text embedding is the bedrock for most Vision-and-Language (...
research
03/28/2020

BiLingUNet: Image Segmentation by Modulating Top-Down and Bottom-Up Visual Processing with Referring Expressions

We present BiLingUNet, a state-of-the-art model for image segmentation u...
research
09/20/2022

Towards Robust Referring Image Segmentation

Referring Image Segmentation (RIS) aims to connect image and language vi...
research
10/24/2022

Towards Unifying Reference Expression Generation and Comprehension

Reference Expression Generation (REG) and Comprehension (REC) are two hi...
research
06/11/2015

Tree-Cut for Probabilistic Image Segmentation

This paper presents a new probabilistic generative model for image segme...
research
11/16/2017

Language-Based Image Editing with Recurrent Attentive Models

We investigate the problem of Language-Based Image Editing (LBIE) in thi...
research
03/06/2013

On Considering Uncertainty and Alternatives in Low-Level Vision

In this paper we address the uncertainty issues involved in the low-leve...

Please sign up or login with your details

Forgot password? Click here to reset