Human-like Controllable Image Captioning with Verb-specific Semantic Roles

03/22/2021
by   Long Chen, et al.
0

Controllable Image Captioning (CIC) – generating image descriptions following designated control signals – has received unprecedented attention over the last few years. To emulate the human ability in controlling caption generation, current CIC studies focus exclusively on control signals concerning objective properties, such as contents of interest or descriptive patterns. However, we argue that almost all existing objective control signals have overlooked two indispensable characteristics of an ideal control signal: 1) Event-compatible: all visual contents referred to in a single sentence should be compatible with the described activity. 2) Sample-suitable: the control signals should be suitable for a specific image sample. To this end, we propose a new control signal for CIC: Verb-specific Semantic Roles (VSR). VSR consists of a verb and some semantic roles, which represents a targeted activity and the roles of entities involved in this activity. Given a designated VSR, we first train a grounded semantic role labeling (GSRL) model to identify and ground all entities for each role. Then, we propose a semantic structure planner (SSP) to learn human-like descriptive semantic structures. Lastly, we use a role-shift captioning model to generate the captions. Extensive experiments and ablations demonstrate that our framework can achieve better controllability than several strong baselines on two challenging CIC benchmarks. Besides, we can generate multi-level diverse captions easily. The code is available at: https://github.com/mad-red/VSR-guided-CIC.

READ FULL TEXT

page 1

page 7

page 8

page 13

research
10/16/2021

Self-Annotated Training for Controllable Image Captioning

The Controllable Image Captioning (CIC) task aims to generate captions c...
research
11/26/2018

Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions

Current captioning approaches can describe images using black-box archit...
research
07/19/2020

Length-Controllable Image Captioning

The last decade has witnessed remarkable progress in the image captionin...
research
12/27/2022

Noise-aware Learning from Web-crawled Image-Text Data for Image Captioning

Image captioning is one of the straightforward tasks that can take advan...
research
03/27/2018

Neural Baby Talk

We introduce a novel framework for image captioning that can produce nat...
research
10/09/2019

Semantic-aware Image Deblurring

Image deblurring has achieved exciting progress in recent years. However...
research
06/21/2022

Multilayer Block Models for Exploratory Analysis of Computer Event Logs

We investigate a graph-based approach to exploratory data analysis in th...

Please sign up or login with your details

Forgot password? Click here to reset