Prompting Large Language Models to Reformulate Queries for Moment Localization

06/06/2023
by   Wenfeng Yan, et al.
0

The task of moment localization is to localize a temporal moment in an untrimmed video for a given natural language query. Since untrimmed video contains highly redundant contents, the quality of the query is crucial for accurately localizing moments, i.e., the query should provide precise information about the target moment so that the localization model can understand what to look for in the videos. However, the natural language queries in current datasets may not be easy to understand for existing models. For example, the Ego4D dataset uses question sentences as the query to describe relatively complex moments. While being natural and straightforward for humans, understanding such question sentences are challenging for mainstream moment localization models like 2D-TAN. Inspired by the recent success of large language models, especially their ability of understanding and generating complex natural language contents, in this extended abstract, we make early attempts at reformulating the moment queries into a set of instructions using large language models and making them more friendly to the localization models.

READ FULL TEXT
research
09/05/2018

Localizing Moments in Video with Temporal Language

Localizing moments in a longer video via natural language queries is a n...
research
05/05/2023

Zelda: Video Analytics using Vision-Language Models

Advances in ML have motivated the design of video analytics systems that...
research
08/10/2022

Exploring Anchor-based Detection for Ego4D Natural Language Query

In this paper we provide the technique report of Ego4D natural language ...
research
10/14/2021

P-Adapters: Robustly Extracting Factual Information from Language Models with Diverse Prompts

Recent work (e.g. LAMA (Petroni et al., 2019)) has found that the qualit...
research
09/22/2021

Natural Language Video Localization with Learnable Moment Proposals

Given an untrimmed video and a natural language query, Natural Language ...
research
04/01/2021

A Survey on Natural Language Video Localization

Natural language video localization (NLVL), which aims to locate a targe...
research
06/05/2023

Overcoming Weak Visual-Textual Alignment for Video Moment Retrieval

Video moment retrieval (VMR) aims to identify the specific moment in an ...

Please sign up or login with your details

Forgot password? Click here to reset