End-to-End Video Question-Answer Generation with Generator-Pretester Network

by   Hung-Ting Su, et al.

We study a novel task, Video Question-Answer Generation (VQAG), for challenging Video Question Answering (Video QA) task in multimedia. Due to expensive data annotation costs, many widely used, large-scale Video QA datasets such as Video-QA, MSVD-QA and MSRVTT-QA are automatically annotated using Caption Question Generation (CapQG) which inputs captions instead of the video itself. As captions neither fully represent a video, nor are they always practically available, it is crucial to generate question-answer pairs based on a video via Video Question-Answer Generation (VQAG). Existing video-to-text (V2T) approaches, despite taking a video as the input, only generate a question alone. In this work, we propose a novel model Generator-Pretester Network that focuses on two components: (1) The Joint Question-Answer Generator (JQAG) which generates a question with its corresponding answer to allow Video Question "Answering" training. (2) The Pretester (PT) verifies a generated question by trying to answer it and checks the pretested answer with both the model's proposed answer and the ground truth answer. We evaluate our system with the only two available large-scale human-annotated Video QA datasets and achieves state-of-the-art question generation performances. Furthermore, using our generated QA pairs only on the Video QA task, we can surpass some supervised baselines. We apply our generated questions to Video QA applications and surpasses some supervised baselines using generated questions only. As a pre-training strategy, we outperform both CapQG and transfer learning approaches when employing semi-supervised (20 with annotated data. These experimental results suggest the novel perspectives for Video QA training.


page 1

page 2

page 8


Video Question Answering with Iterative Video-Text Co-Tokenization

Video question answering is a challenging task that requires understandi...

Addressing Semantic Drift in Question Generation for Semi-Supervised Question Answering

Text-based Question Generation (QG) aims at generating natural and relev...

Self-supervised pre-training and contrastive representation learning for multiple-choice video QA

Video Question Answering (Video QA) requires fine-grained understanding ...

Leveraging Video Descriptions to Learn Video Question Answering

We propose a scalable approach to learn video-based question answering (...

Weakly Supervised Visual Question Answer Generation

Growing interest in conversational agents promote twoway human-computer ...

Dataset and Neural Recurrent Sequence Labeling Model for Open-Domain Factoid Question Answering

While question answering (QA) with neural network, i.e. neural QA, has a...

Generating Diverse and Consistent QA pairs from Contexts with Information-Maximizing Hierarchical Conditional VAEs

One of the most crucial challenges in questionanswering (QA) is the scar...

Please sign up or login with your details

Forgot password? Click here to reset