An Overview of Cross-media Retrieval: Concepts, Methodologies, Benchmarks and Challenges

by   Yuxin Peng, et al.

Multimedia retrieval plays an indispensable role in big data utilization. Past efforts mainly focused on single-media retrieval. However, the requirements of users are highly flexible, such as retrieving the relevant audio clips with one query of image. So challenges stemming from the "media gap", which means that representations of different media types are inconsistent, have attracted increasing attention. Cross-media retrieval is designed for the scenarios where the queries and retrieval results are of different media types. As a relatively new research topic, its concepts, methodologies and benchmarks are still not clear in the literatures. To address these issues, we review more than 100 references, give an overview including the concepts, methodologies, major challenges and open issues, as well as build up the benchmarks including datasets and experimental results. Researchers can directly adopt the benchmarks to promptly evaluate their proposed methods. This will help them to focus on algorithm design, rather than the time-consuming compared methods and results. It is noted that we have constructed a new dataset XMedia, which is the first publicly available dataset with up to five media types (text, image, video, audio and 3D model). We believe this overview will attract more researchers to focus on cross-media retrieval and be helpful to them.


page 1

page 2

page 14


A New Benchmark and Approach for Fine-grained Cross-media Retrieval

Cross-media retrieval is to return the results of various media types co...

Deep Learning Techniques for Future Intelligent Cross-Media Retrieval

With the advancement in technology and the expansion of broadcasting, cr...

Unsupervised Cross-Media Hashing with Structure Preservation

Recent years have seen the exponential growth of heterogeneous multimedi...

Deep Cross-media Knowledge Transfer

Cross-media retrieval is a research hotspot in multimedia area, which ai...

Cross-media Multi-level Alignment with Relation Attention Network

With the rapid growth of multimedia data, such as image and text, it is ...

Benchmarks, Performance Evaluation and Contests for 3D Shape Retrieval

Benchmarking of 3D Shape retrieval allows developers and researchers to ...

Deepfake: Definitions, Performance Metrics and Standards, Datasets and Benchmarks, and a Meta-Review

Recent advancements in AI, especially deep learning, have contributed to...

Please sign up or login with your details

Forgot password? Click here to reset