SuperPAL: Supervised Proposition ALignment for Multi-Document Summarization and Derivative Sub-Tasks

09/01/2020
by   Ori Ernst, et al.
0

Multi-document summarization (MDS) is a challenging task, often decomposed to subtasks of salience and redundancy detection, followed by generation. While alignment of spans between reference summaries and source documents has been leveraged for training component tasks, the underlying alignment step was never independently addressed or evaluated. We advocate developing high quality source-reference alignment algorithms, that can be applied to recent large-scale datasets to obtain useful "silver", i.e. approximate, training data. As a first step, we present an annotation methodology by which we create gold standard development and test sets for summary-source alignment, and suggest its utility for tuning and evaluating effective alignment algorithms, as well as for properly evaluating MDS subtasks. Second, we introduce a new large-scale alignment dataset for training, with which an automatic alignment model was trained. This aligner achieves higher coherency with the reference summary than previous aligners used for summarization, and gets significantly higher ROUGE results when replacing a simpler aligner in a competitive summarization model. Finally, we release three additional datasets (for salience, clustering and generation), naturally derived from our alignment datasets. Furthermore, these datasets can be derived from any summarization dataset automatically after extracting alignments with our trained aligner. Hence, they can be utilized for training summarization sub-tasks.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
04/24/2018

Towards a Neural Network Approach to Abstractive Multi-Document Summarization

Till now, neural abstractive summarization methods have achieved great s...
research
11/03/2020

WSL-DS: Weakly Supervised Learning with Distant Supervision for Query Focused Multi-Document Abstractive Summarization

In the Query Focused Multi-Document Summarization (QF-MDS) task, a set o...
research
05/11/2022

ALIGNMEET: A Comprehensive Tool for Meeting Annotation, Alignment, and Evaluation

Summarization is a challenging problem, and even more challenging is to ...
research
04/30/2020

TLDR: Extreme Summarization of Scientific Documents

We introduce TLDR generation for scientific papers, a new automatic summ...
research
10/07/2021

HowSumm: A Multi-Document Summarization Dataset Derived from WikiHow Articles

We present HowSumm, a novel large-scale dataset for the task of query-fo...
research
12/16/2021

A Proposition-Level Clustering Approach for Multi-Document Summarization

Text clustering methods were traditionally incorporated into multi-docum...
research
11/26/2015

TGSum: Build Tweet Guided Multi-Document Summarization Dataset

The development of summarization research has been significantly hampere...

Please sign up or login with your details

Forgot password? Click here to reset