Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning

11/21/2017
by   Qi Wu, et al.
0

The Visual Dialogue task requires an agent to engage in a conversation about an image with a human. It represents an extension of the Visual Question Answering task in that the agent needs to answer a question about an image, but it needs to do so in light of the previous dialogue that has taken place. The key challenge in Visual Dialogue is thus maintaining a consistent, and natural dialogue while continuing to answer questions correctly. We present a novel approach that combines Reinforcement Learning and Generative Adversarial Networks (GANs) to generate more human-like responses to questions. The GAN helps overcome the relative paucity of training data, and the tendency of the typical MLE-based approach to generate overly terse answers. Critically, the GAN is tightly integrated into the attention mechanism that generates human-interpretable reasons for each answer. This means that the discriminative model of the GAN has the task of assessing whether a candidate answer is generated by a human or not, given the provided reason. This is significant because it drives the generative model to produce high quality answers that are well supported by the associated reasoning. The method also generates the state-of-the-art results on the primary benchmark.

READ FULL TEXT
research
02/11/2018

FlipDial: A Generative Model for Two-Way Visual Dialogue

We present FlipDial, a generative model for visual dialogue that simulta...
research
11/17/2019

DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue

Different from Visual Question Answering task that requires to answer on...
research
01/23/2017

Adversarial Learning for Neural Dialogue Generation

In this paper, drawing intuition from the Turing test, we propose using ...
research
08/12/2019

Why Does a Visual Question Have Different Answers?

Visual question answering is the task of returning the answer to a quest...
research
04/20/2020

A Revised Generative Evaluation of Visual Dialogue

Evaluating Visual Dialogue, the task of answering a sequence of question...
research
11/06/2019

A Spoken Dialogue System for Spatial Question Answering in a Physical Blocks World

The blocks world is a classic toy domain that has long been used to buil...
research
10/01/2020

Answer-Driven Visual State Estimator for Goal-Oriented Visual Dialogue

A goal-oriented visual dialogue involves multi-turn interactions between...

Please sign up or login with your details

Forgot password? Click here to reset