GameWikiSum: a Novel Large Multi-Document Summarization Dataset

02/17/2020
by   Diego Antognini, et al.
0

Today's research progress in the field of multi-document summarization is obstructed by the small number of available datasets. Since the acquisition of reference summaries is costly, existing datasets contain only hundreds of samples at most, resulting in heavy reliance on hand-crafted features or necessitating additional, manually annotated data. The lack of large corpora therefore hinders the development of sophisticated models. Additionally, most publicly available multi-document summarization corpora are in the news domain, and no analogous dataset exists in the video game domain. In this paper, we propose GameWikiSum, a new domain-specific dataset for multi-document summarization, which is one hundred times larger than commonly used datasets, and in another domain than news. Input documents consist of long professional video game reviews as well as references of their gameplay sections in Wikipedia pages. We analyze the proposed dataset and show that both abstractive and extractive models can be trained on it. We release GameWikiSum for further research: https://github.com/Diego999/GameWikiSum.

READ FULL TEXT
research
10/23/2020

AQuaMuSe: Automatically Generating Datasets for Query-Based Multi-Document Summarization

Summarization is the task of compressing source document(s) into coheren...
research
04/13/2021

MS2: Multi-Document Summarization of Medical Studies

To assess the effectiveness of any medical intervention, researchers mus...
research
07/18/2022

GOAL: Towards Benchmarking Few-Shot Sports Game Summarization

Sports game summarization aims to generate sports news based on real-tim...
research
11/15/2020

Open4Business(O4B): An Open Access Dataset for Summarizing Business Documents

A major challenge in fine-tuning deep learning models for automatic summ...
research
09/16/2023

ODSum: New Benchmarks for Open Domain Multi-Document Summarization

Open-domain Multi-Document Summarization (ODMDS) is a critical tool for ...
research
05/27/2023

MeetingBank: A Benchmark Dataset for Meeting Summarization

As the number of recorded meetings increases, it becomes increasingly im...
research
11/12/2018

CQASUMM: Building References for Community Question Answering Summarization Corpora

Community Question Answering forums such as Quora, Stackoverflow are ric...

Please sign up or login with your details

Forgot password? Click here to reset