Structured Knowledge Discovery from Massive Text Corpus

07/23/2019
by   Chenwei Zhang, et al.
0

Nowadays, with the booming development of the Internet, people benefit from its convenience due to its open and sharing nature. A large volume of natural language texts is being generated by users in various forms, such as search queries, documents, and social media posts. As the unstructured text corpus is usually noisy and messy, it becomes imperative to correctly identify and accurately annotate structured information in order to obtain meaningful insights or better understand unstructured texts. On the other hand, the existing structured information, which embodies our knowledge such as entity or concept relations, often suffers from incompleteness or quality-related issues. Given a gigantic collection of texts which offers rich semantic information, it is also important to harness the massiveness of the unannotated text corpus to expand and refine existing structured knowledge with fewer annotation efforts. In this dissertation, I will introduce principles, models, and algorithms for effective structured knowledge discovery from the massive text corpus. We are generally interested in obtaining insights and better understanding unstructured texts with the help of structured annotations or by structure-aware modeling. Also, given the existing structured knowledge, we are interested in expanding its scale and improving its quality harnessing the massiveness of the text corpus. In particular, four problems are studied in this dissertation: Structured Intent Detection for Natural Language Understanding, Structure-aware Natural Language Modeling, Generative Structured Knowledge Expansion, and Synonym Refinement on Structured Knowledge.

READ FULL TEXT

page 1

page 2

page 3

page 4

03/15/2017

InScript: Narrative texts annotated with script information

This paper presents the InScript corpus (Narrative Texts Instantiating S...
12/28/2021

Cognitive Computing to Optimize IT Services

In this paper, the challenges of maintaining a healthy IT operational en...
11/30/2015

Ask, and shall you receive?: Understanding Desire Fulfillment in Natural Language Text

The ability to comprehend wishes or desires and their fulfillment is imp...
06/21/2019

SurfCon: Synonym Discovery on Privacy-Aware Clinical Data

Unstructured clinical texts contain rich health-related information. To ...
12/04/2019

Implicit Knowledge in Argumentative Texts: An Annotated Corpus

When speaking or writing, people omit information that seems clear and e...
06/02/2021

Generating Informative Conclusions for Argumentative Texts

The purpose of an argumentative text is to support a certain conclusion....
02/04/2020

Plague Dot Text: Text mining and annotation of outbreak reports of the Third Plague Pandemic (1894-1952)

The design of models that govern diseases in population is commonly buil...