A survey on knowledge-enhanced multimodal learning

11/19/2022
by   Maria Lymperaiou, et al.
0

Multimodal learning has been a field of increasing interest, aiming to combine various modalities in a single joint representation. Especially in the area of visiolinguistic (VL) learning multiple models and techniques have been developed, targeting a variety of tasks that involve images and text. VL models have reached unprecedented performances by extending the idea of Transformers, so that both modalities can learn from each other. Massive pre-training procedures enable VL models to acquire a certain level of real-world understanding, although many gaps can be identified: the limited comprehension of commonsense, factual, temporal and other everyday knowledge aspects questions the extendability of VL tasks. Knowledge graphs and other knowledge sources can fill those gaps by explicitly providing missing information, unlocking novel capabilities of VL models. In the same time, knowledge graphs enhance explainability, fairness and validity of decision making, issues of outermost importance for such complex implementations. The current survey aims to unify the fields of VL representation learning and knowledge graphs, and provides a taxonomy and analysis of knowledge-enhanced VL models.

READ FULL TEXT

page 22

page 23

research
03/04/2023

The Contribution of Knowledge in Visiolinguistic Learning: A Survey on Tasks and Challenges

Recent advancements in visiolinguistic (VL) learning have allowed the de...
research
02/19/2020

Error detection in Knowledge Graphs: Path Ranking, Embeddings or both?

This paper attempts to compare and combine different approaches for de-t...
research
05/27/2019

Relational Representation Learning for Dynamic (Knowledge) Graphs: A Survey

Graphs arise naturally in many real-world applications including social ...
research
09/11/2021

Discovering Technology Gaps using the IntSight Knowledge Navigator

Knowledge analysis is an important application of knowledge graphs. In t...
research
06/06/2023

MolFM: A Multimodal Molecular Foundation Model

Molecular knowledge resides within three different modalities of informa...
research
08/27/2020

A Taxonomy of Knowledge Gaps for Wikimedia Projects (First Draft)

In January 2019, prompted by the Wikimedia Movement's 2030 strategic dir...
research
07/14/2017

Knowledge will Propel Machine Understanding of Content: Extrapolating from Current Examples

Machine Learning has been a big success story during the AI resurgence. ...

Please sign up or login with your details

Forgot password? Click here to reset