WikiGraphs: A Wikipedia Text - Knowledge Graph Paired Dataset

07/20/2021
by   Luyu Wang, et al.
65

We present a new dataset of Wikipedia articles each paired with a knowledge graph, to facilitate the research in conditional text generation, graph generation and graph representation learning. Existing graph-text paired datasets typically contain small graphs and short text (1 or few sentences), thus limiting the capabilities of the models that can be learned on the data. Our new dataset WikiGraphs is collected by pairing each Wikipedia article from the established WikiText-103 benchmark (Merity et al., 2016) with a subgraph from the Freebase knowledge graph (Bollacker et al., 2008). This makes it easy to benchmark against other state-of-the-art text generative models that are capable of generating long paragraphs of coherent text. Both the graphs and the text data are of significantly larger scale compared to prior graph-text paired datasets. We present baseline graph neural network and transformer model results on our dataset for 3 tasks: graph -> text generation, graph -> text retrieval and text -> graph retrieval. We show that better conditioning on the graph provides gains in generation and retrieval quality but there is still large room for improvement.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/30/2021

EventNarrative: A large-scale Event-centric Dataset for Knowledge Graph-to-Text Generation

We introduce EventNarrative, a knowledge graph-to-text dataset from publ...
research
07/04/2023

Knowledge Graph for NLG in the context of conversational agents

The use of knowledge graphs (KGs) enhances the accuracy and comprehensiv...
research
12/29/2020

Generating Wikipedia Article Sections from Diverse Data Sources

Datasets for data-to-text generation typically focus either on multi-dom...
research
09/20/2023

Construction of Paired Knowledge Graph-Text Datasets Informed by Cyclic Evaluation

Datasets that pair Knowledge Graphs (KG) and text together (KG-T) can be...
research
06/15/2022

KGEA: A Knowledge Graph Enhanced Article Quality Identification Dataset

With so many articles of varying quality being produced at every moment,...
research
06/02/2020

Graph-Stega: Semantic Controllable Steganographic Text Generation Guided by Knowledge Graph

Most of the existing text generative steganographic methods are based on...
research
12/16/2021

FRUIT: Faithfully Reflecting Updated Information in Text

Textual knowledge bases such as Wikipedia require considerable effort to...

Please sign up or login with your details

Forgot password? Click here to reset