MHSCNet: A Multimodal Hierarchical Shot-aware Convolutional Network for Video Summarization

04/18/2022
by   Wujiang Xu, et al.
0

Video summarization intends to produce a concise video summary by effectively capturing and combining the most informative parts of the whole content. Existing approaches for video summarization regard the task as a frame-wise keyframe selection problem and generally construct the frame-wise representation by combining the long-range temporal dependency with the unimodal or bimodal information. However, the optimal video summaries need to reflect the most valuable keyframe with its own information, and one with semantic power of the whole content. Thus, it is critical to construct a more powerful and robust frame-wise representation and predict the frame-level importance score in a fair and comprehensive manner. To tackle the above issues, we propose a multimodal hierarchical shot-aware convolutional network, denoted as MHSCNet, to enhance the frame-wise representation via combining the comprehensive available multimodal information. Specifically, we design a hierarchical ShotConv network to incorporate the adaptive shot-aware frame-level representation by considering the short-range and long-range temporal dependency. Based on the learned shot-aware representations, MHSCNet can predict the frame-level importance score in the local and global view of the video. Extensive experiments on two standard video summarization datasets demonstrate that our proposed method consistently outperforms state-of-the-art baselines. Source code will be made publicly available.

READ FULL TEXT

page 4

page 7

page 8

research
05/10/2021

Reconstructive Sequence-Graph Network for Video Summarization

Exploiting the inner-shot and inter-shot dependencies is essential for k...
research
01/26/2023

LoRaLay: A Multilingual and Multimodal Dataset for Long Range and Layout-Aware Summarization

Text Summarization is a popular task and an active area of research for ...
research
08/23/2017

CNN-Based Prediction of Frame-Level Shot Importance for Video Summarization

In the Internet, ubiquitous presence of redundant, unedited, raw videos ...
research
06/13/2023

360TripleView: 360-Degree Video View Management System Driven by Convergence Value of Viewing Preferences

360-degree video has become increasingly popular in content consumption....
research
01/30/2015

Co-Regularized Deep Representations for Video Summarization

Compact keyframe-based video summaries are a popular way of generating v...
research
11/24/2018

Discriminative Feature Learning for Unsupervised Video Summarization

In this paper, we address the problem of unsupervised video summarizatio...
research
09/08/2021

VideoModerator: A Risk-aware Framework for Multimodal Video Moderation in E-Commerce

Video moderation, which refers to remove deviant or explicit content fro...

Please sign up or login with your details

Forgot password? Click here to reset