Image Memorability Prediction with Vision Transformers

01/20/2023
by   Thomas Hagen, et al.
0

Behavioral studies have shown that the memorability of images is similar across groups of people, suggesting that memorability is a function of the intrinsic properties of images, and is unrelated to people's individual experiences and traits. Deep learning networks can be trained on such properties and be used to predict memorability in new data sets. Convolutional neural networks (CNN) have pioneered image memorability prediction, but more recently developed vision transformer (ViT) models may have the potential to yield even better predictions. In this paper, we present the ViTMem, a new memorability model based on ViT, and evaluate memorability predictions obtained by it with state-of-the-art CNN-derived models. Results showed that ViTMem performed equal to or better than state-of-the-art models on all data sets. Additional semantic level analyses revealed that ViTMem is particularly sensitive to the semantic content that drives memorability in images. We conclude that ViTMem provides a new step forward, and propose that ViT-derived models can replace CNNs for computational prediction of image memorability. Researchers, educators, advertisers, visual designers and other interested parties can leverage the model to improve the memorability of their image material.

READ FULL TEXT
research
05/21/2021

Embracing New Techniques in Deep Learning for Estimating Image Memorability

Various work has suggested that the memorability of an image is consiste...
research
06/01/2022

A comparative study between vision transformers and CNNs in digital pathology

Recently, vision transformers were shown to be capable of outperforming ...
research
07/17/2020

How Flexible is that Functional Form? Quantifying the Restrictiveness of Theories

We propose a new way to quantify the restrictiveness of an economic mode...
research
03/08/2019

Image Privacy Prediction Using Deep Neural Networks

Images today are increasingly shared online on social networking sites s...
research
04/30/2021

Interpretable Semantic Photo Geolocalization

Planet-scale photo geolocalization is the complex task of estimating the...
research
10/14/2021

Interactive Analysis of CNN Robustness

While convolutional neural networks (CNNs) have found wide adoption as s...

Please sign up or login with your details

Forgot password? Click here to reset