Resolving References to Objects in Photographs using the Words-As-Classifiers Model

10/07/2015
by   David Schlangen, et al.
0

A common use of language is to refer to visually present objects. Modelling it in computers requires modelling the link between language and perception. The "words as classifiers" model of grounded semantics views words as classifiers of perceptual contexts, and composes the meaning of a phrase through composition of the denotations of its component words. It was recently shown to perform well in a game-playing scenario with a small number of object types. We apply it to two large sets of real-world photographs that contain a much larger variety of types and for which referring expressions are available. Using a pre-trained convolutional neural network to extract image features, and augmenting these with in-picture positional information, we show that the model achieves performance competitive with the state of the art in a reference resolution task (given expression, find bounding box of its referent), while, as we argue, being conceptually simpler and more flexible.

READ FULL TEXT

page 4

page 5

research
11/08/2019

Composing and Embedding the Words-as-Classifiers Model of Grounded Semantics

The words-as-classifiers model of grounded lexical semantics learns a se...
research
06/12/2018

iParaphrasing: Extracting Visually Grounded Paraphrases via an Image

A paraphrase is a restatement of the meaning of a text in other words. P...
research
04/04/2023

Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQA

In this paper, we propose a novel multi-modal framework for Scene Text V...
research
12/17/2018

Grounded Video Description

Video description is one of the most challenging problems in vision and ...
research
04/10/2017

Pay Attention to Those Sets! Learning Quantification from Images

Major advances have recently been made in merging language and vision re...
research
06/28/2016

"Show me the cup": Reference with Continuous Representations

One of the most basic functions of language is to refer to objects in a ...
research
02/25/2023

Vagueness in Predicates and Objects

Classical semantics assumes that one can model reference, predication an...

Please sign up or login with your details

Forgot password? Click here to reset