PDSum: Prototype-driven Continuous Summarization of Evolving Multi-document Sets Stream

02/10/2023
by   Susik Yoon, et al.
0

Summarizing text-rich documents has been long studied in the literature, but most of the existing efforts have been made to summarize a static and predefined multi-document set. With the rapid development of online platforms for generating and distributing text-rich documents, there arises an urgent need for continuously summarizing dynamically evolving multi-document sets where the composition of documents and sets is changing over time. This is especially challenging as the summarization should be not only effective in incorporating relevant, novel, and distinctive information from each concurrent multi-document set, but also efficient in serving online applications. In this work, we propose a new summarization problem, Evolving Multi-Document sets stream Summarization (EMDS), and introduce a novel unsupervised algorithm PDSum with the idea of prototype-driven continuous summarization. PDSum builds a lightweight prototype of each multi-document set and exploits it to adapt to new documents while preserving accumulated knowledge from previous documents. To update new summaries, the most representative sentences for each multi-document set are extracted by measuring their similarities to the prototypes. A thorough evaluation with real multi-document sets streams demonstrates that PDSum outperforms state-of-the-art unsupervised multi-document summarization algorithms in EMDS in terms of relevance, novelty, and distinctiveness and is also robust to various evaluation settings.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/05/2023

Mining both Commonality and Specificity from Multiple Documents for Multi-Document Summarization

The multi-document summarization task requires the designed summarizer t...
research
04/11/2023

LBMT team at VLSP2022-Abmusu: Hybrid method with text correlation and generative models for Vietnamese multi-document summarization

Multi-document summarization is challenging because the summaries should...
research
08/19/2018

Adapting the Neural Encoder-Decoder Framework from Single to Multi-Document Summarization

Generating an abstract from a set of relevant documents remains challeng...
research
10/06/2020

SupMMD: A Sentence Importance Model for Extractive Summarization using Maximum Mean Discrepancy

Most work on multi-document summarization has focused on generic summari...
research
10/03/2021

Multi-Document Keyphrase Extraction: A Literature Review and the First Dataset

Keyphrase extraction has been comprehensively researched within the sing...
research
05/12/2016

Real-Time Web Scale Event Summarization Using Sequential Decision Making

We present a system based on sequential decision making for the online s...
research
08/06/2015

Privacy-Preserving Multi-Document Summarization

State-of-the-art extractive multi-document summarization systems are usu...

Please sign up or login with your details

Forgot password? Click here to reset