Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

02/28/2022
by   Luís Vilaça, et al.
0

Audio-visual correlation learning aims to capture essential correspondences and understand natural phenomena between audio and video. With the rapid growth of deep learning, an increasing amount of attention has been paid to this emerging research issue. Through the past few years, various methods and datasets have been proposed for audio-visual correlation learning, which motivate us to conclude a comprehensive survey. This survey paper focuses on state-of-the-art (SOTA) models used to learn correlations between audio and video, but also discusses some tasks of definition and paradigm applied in AI multimedia. In addition, we investigate some objective functions frequently used for optimizing audio-visual correlation learning models and discuss how audio-visual data is exploited in the optimization process. Most importantly, we provide an extensive comparison and summarization of the recent progress of SOTA audio-visual correlation learning and discuss future research directions.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
12/02/2022

Role of Audio in Audio-Visual Video Summarization

Video summarization attracts attention for efficient video representatio...
research
08/08/2022

Abstractive Meeting Summarization: A Survey

Recent advances in deep learning, and especially the invention of encode...
research
01/14/2020

Deep Audio-Visual Learning: A Survey

Audio-visual learning, aimed at exploiting the relationship between audi...
research
07/20/2021

Data Hiding with Deep Learning: A Survey Unifying Digital Watermarking and Steganography

Data hiding is the process of embedding information into a noise-toleran...
research
07/10/2023

A Demand-Driven Perspective on Generative Audio AI

To achieve successful deployment of AI research, it is crucial to unders...
research
08/20/2022

Learning in Audio-visual Context: A Review, Analysis, and New Perspective

Sight and hearing are two senses that play a vital role in human communi...
research
06/20/2022

A Comprehensive Survey on Video Saliency Detection with Auditory Information: the Audio-visual Consistency Perceptual is the Key!

Video saliency detection (VSD) aims at fast locating the most attractive...

Please sign up or login with your details

Forgot password? Click here to reset