Cascaded Mutual Modulation for Visual Reasoning

09/06/2018
by   Yiqun Yao, et al.
0

Visual reasoning is a special visual question answering problem that is multi-step and compositional by nature, and also requires intensive text-vision interactions. We propose CMM: Cascaded Mutual Modulation as a novel end-to-end visual reasoning model. CMM includes a multi-step comprehension process for both question and image. In each step, we use a Feature-wise Linear Modulation (FiLM) technique to enable textual/visual pipeline to mutually control each other. Experiments show that CMM significantly outperforms most related models, and reach state-of-the-arts on two visual reasoning benchmarks: CLEVR and NLVR, collected from both synthetic and natural languages. Ablation studies confirm that both our multistep framework and our visual-guided language modulation are critical to the task. Our code is available at https://github.com/FlamingHorizon/CMM-VR.

READ FULL TEXT
research
11/21/2019

Temporal Reasoning via Audio Question Answering

Multimodal question answering tasks can be used as proxy tasks to study ...
research
07/27/2020

REXUP: I REason, I EXtract, I UPdate with Structured Compositional Reasoning for Visual Question Answering

Visual question answering (VQA) is a challenging multi-modal task that r...
research
05/06/2022

QLEVR: A Diagnostic Dataset for Quantificational Language and Elementary Visual Reasoning

Synthetic datasets have successfully been used to probe visual question-...
research
11/26/2020

Transformation Driven Visual Reasoning

This paper defines a new visual reasoning paradigm by introducing an imp...
research
09/22/2017

FiLM: Visual Reasoning with a General Conditioning Layer

We introduce a general-purpose conditioning method for neural networks c...
research
01/03/2019

CLEVR-Ref+: Diagnosing Visual Reasoning with Referring Expressions

Referring object detection and referring image segmentation are importan...
research
03/26/2022

Visual Abductive Reasoning

Abductive reasoning seeks the likeliest possible explanation for partial...

Please sign up or login with your details

Forgot password? Click here to reset