UNION: An Unreferenced Metric for Evaluating Open-ended Story Generation

09/16/2020
by   Jian Guan, et al.
0

Despite the success of existing referenced metrics (e.g., BLEU and MoverScore), they correlate poorly with human judgments for open-ended text generation including story or dialog generation because of the notorious one-to-many issue: there are many plausible outputs for the same input, which may differ substantially in literal or semantics from the limited number of given references. To alleviate this issue, we propose UNION, a learnable unreferenced metric for evaluating open-ended story generation, which measures the quality of a generated story without any reference. Built on top of BERT, UNION is trained to distinguish human-written stories from negative samples and recover the perturbation in negative stories. We propose an approach of constructing negative samples by mimicking the errors commonly observed in existing NLG models, including repeated plots, conflicting logic, and long-range incoherence. Experiments on two story datasets demonstrate that UNION is a reliable measure for evaluating the quality of generated stories, which correlates better with human judgments and is more generalizable than existing state-of-the-art metrics.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/19/2021

OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics

Automatic metrics are essential for developing natural language generati...
research
04/12/2021

Plot-guided Adversarial Example Construction for Evaluating Open-domain Story Generation

With the recent advances of open-domain story generation, the lack of re...
research
03/15/2023

DeltaScore: Evaluating Story Generation with Differentiating Perturbations

Various evaluation metrics exist for natural language generation tasks, ...
research
09/14/2021

The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation

Recent text generation research has increasingly focused on open-ended d...
research
05/13/2018

Hierarchical Neural Story Generation

We explore story generation: creative systems that can build coherent an...
research
08/22/2023

StoryBench: A Multifaceted Benchmark for Continuous Story Visualization

Generating video stories from text prompts is a complex task. In additio...
research
09/11/2019

What Makes A Good Story? Designing Composite Rewards for Visual Storytelling

Previous storytelling approaches mostly focused on optimizing traditiona...

Please sign up or login with your details

Forgot password? Click here to reset