DiscoFuse: A Large-Scale Dataset for Discourse-based Sentence Fusion

02/27/2019
by   Mor Geva, et al.
0

Sentence fusion is the task of joining several independent sentences into a single coherent text. Current datasets for sentence fusion are small and insufficient for training modern neural models. In this paper, we propose a method for automatically-generating fusion examples from raw text and present DiscoFuse, a large scale dataset for discourse-based sentence fusion. We author a set of rules for identifying a diverse set of discourse phenomena in raw text, and decomposing the text into two independent sentences. We apply our approach on two document collections: Wikipedia and Sports articles, yielding 60 million fusion examples annotated with discourse information required to reconstruct the fused text. We develop a sequence-to-sequence model on DiscoFuse and thoroughly analyze its strengths and weaknesses with respect to the various discourse phenomena, using both automatic as well as human evaluation. Finally, we conduct transfer learning experiments with WebSplit, a recent dataset for text simplification. We show that pretraining on DiscoFuse substantially improves performance on WebSplit when viewed as a sentence fusion task.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/06/2020

Semantically Driven Sentence Fusion: Modeling and Evaluation

Sentence fusion is the task of joining related sentences into coherent t...
research
03/28/2019

Mining Discourse Markers for Unsupervised Sentence Representation Learning

Current state of the art systems in NLP heavily rely on manually annotat...
research
09/05/2019

TransSent: Towards Generation of Structured Sentences with Discourse Marker

This paper focuses on the task of generating long structured sentences w...
research
10/11/2021

Document-Level Text Simplification: Dataset, Criteria and Baseline

Text simplification is a valuable technique. However, current research i...
research
09/18/2022

Improving Topic Segmentation by Injecting Discourse Dependencies

Recent neural supervised topic segmentation models achieve distinguished...
research
10/30/2022

Actionable Phrase Detection using NLP

Actionable sentences are terms that, in the most basic sense, imply the ...
research
02/03/2017

Automatic Prediction of Discourse Connectives

Accurate prediction of suitable discourse connectives (however, furtherm...

Please sign up or login with your details

Forgot password? Click here to reset