Recipe Generation from Unsegmented Cooking Videos

09/21/2022
by   Taichi Nishimura, et al.
0

This paper tackles recipe generation from unsegmented cooking videos, a task that requires agents to (1) extract key events in completing the dish and (2) generate sentences for the extracted events. Our task is similar to dense video captioning (DVC), which aims at detecting events thoroughly and generating sentences for them. However, unlike DVC, in recipe generation, recipe story awareness is crucial, and a model should output an appropriate number of key events in the correct order. We analyze the output of the DVC model and observe that although (1) several events are adoptable as a recipe story, (2) the generated sentences for such events are not grounded in the visual content. Based on this, we hypothesize that we can obtain correct recipes by selecting oracle events from the output events of the DVC model and re-generating sentences for them. To achieve this, we propose a novel transformer-based joint approach of training event selector and sentence generator for selecting oracle events from the outputs of the DVC model and generating grounded sentences for the events, respectively. In addition, we extend the model by including ingredients to generate more accurate recipes. The experimental results show that the proposed method outperforms state-of-the-art DVC models. We also confirm that, by modeling the recipe in a story-aware manner, the proposed model output the appropriate number of events in the correct order.

READ FULL TEXT

page 1

page 6

page 9

research
02/05/2021

GraphPlan: Story Generation by Planning with Event Graph

Story generation is a task that aims to automatically produce multiple s...
research
09/08/2019

Story Realization: Expanding Plot Events into Sentences

Neural network based approaches to automated story plot generation attem...
research
06/05/2017

Event Representations for Automated Story Generation with Deep Neural Nets

Automated story generation is the problem of automatically selecting a s...
research
06/24/2020

Comprehensive Information Integration Modeling Framework for Video Titling

In e-commerce, consumer-generated videos, which in general deliver consu...
research
10/16/2012

Semantic Understanding of Professional Soccer Commentaries

This paper presents a novel approach to the problem of semantic parsing ...
research
06/14/2019

"My Way of Telling a Story": Persona based Grounded Story Generation

Visual storytelling is the task of generating stories based on a sequenc...
research
05/31/2020

"Judge me by my size (noun), do you?” YodaLib: A Demographic-Aware Humor Generation Framework

The subjective nature of humor makes computerized humor generation a cha...

Please sign up or login with your details

Forgot password? Click here to reset