AlignVE: Visual Entailment Recognition Based on Alignment Relations

11/16/2022
by   Biwei Cao, et al.
0

Visual entailment (VE) is to recognize whether the semantics of a hypothesis text can be inferred from the given premise image, which is one special task among recent emerged vision and language understanding tasks. Currently, most of the existing VE approaches are derived from the methods of visual question answering. They recognize visual entailment by quantifying the similarity between the hypothesis and premise in the content semantic features from multi modalities. Such approaches, however, ignore the VE's unique nature of relation inference between the premise and hypothesis. Therefore, in this paper, a new architecture called AlignVE is proposed to solve the visual entailment problem with a relation interaction method. It models the relation between the premise and hypothesis as an alignment matrix. Then it introduces a pooling operation to get feature vectors with a fixed size. Finally, it goes through the fully-connected layer and normalization layer to complete the classification. Experiments show that our alignment-based architecture reaches 72.45% accuracy on SNLI-VE dataset, outperforming previous content-based models under the same settings.

READ FULL TEXT

page 2

page 7

page 10

research
11/26/2018

Visual Entailment Task for Visually-Grounded Language Learning

We introduce a new inference task - Visual Entailment (VE) - which diffe...
research
04/16/2021

Multivalent Entailment Graphs for Question Answering

Drawing inferences between open-domain natural language predicates is a ...
research
07/05/2019

A Study of the Effect of Resolving Negation and Sentiment Analysis in Recognizing Text Entailment for Arabic

Recognizing the entailment relation showed that its influence to extract...
research
06/25/2021

Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training

Vision-Language Pre-training (VLP) aims to learn multi-modal representat...
research
01/31/2014

Experiments with Three Approaches to Recognizing Lexical Entailment

Inference in natural language often involves recognizing lexical entailm...
research
07/23/2022

Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language Explanations

Visual Entailment with natural language explanations aims to infer the r...
research
05/03/2020

Bayesian Entailment Hypothesis: How Brains Implement Monotonic and Non-monotonic Reasoning

Recent success of Bayesian methods in neuroscience and artificial intell...

Please sign up or login with your details

Forgot password? Click here to reset