Face-Cap: Image Captioning using Facial Expression Analysis

07/06/2018
by   Omid Mohamad Nezami, et al.
0

Image captioning is the process of generating a natural language description of an image. Most current image captioning models, however, do not take into account the emotional aspect of an image, which is very relevant to activities and interpersonal relationships represented therein. Towards developing a model that can produce human-like captions incorporating these, we use facial expression features extracted from images including human faces, with the aim of improving the descriptive ability of the model. In this work, we present two variants of our Face-Cap model, which embed facial expression features in different ways, to generate image captions. Using all standard evaluation metrics, our Face-Cap models outperform a state-of-the-art baseline model for generating image captions when applied to an image caption dataset extracted from the standard Flickr 30K dataset, consisting of around 11K images containing faces. An analysis of the captions finds that, perhaps surprisingly, the improvement in caption quality appears to come not from the addition of adjectives linked to emotional aspects of the images, but from more variety in the actions described in the captions.

READ FULL TEXT

page 4

page 13

research
08/08/2019

Image Captioning using Facial Expression and Attention

Benefiting from advances in machine vision and natural language processi...
research
01/30/2018

Image Captioning at Will: A Versatile Scheme for Effectively Injecting Sentiments into Image Descriptions

Automatic image captioning has recently approached human-level performan...
research
07/22/2020

Integrating Image Captioning with Rule-based Entity Masking

Given an image, generating its natural language description (i.e., capti...
research
02/06/2018

Multimodal Image Captioning for Marketing Analysis

Automatically captioning images with natural language sentences is an im...
research
03/12/2018

Discriminability objective for training descriptive captions

One property that remains lacking in image captions generated by contemp...
research
04/25/2020

How to read faces without looking at them

Face reading is the most intuitive aspect of emotion recognition. Unfort...
research
06/05/2023

Cheap-fake Detection with LLM using Prompt Engineering

The misuse of real photographs with conflicting image captions in news i...

Please sign up or login with your details

Forgot password? Click here to reset