Not just about size - A Study on the Role of Distributed Word Representations in the Analysis of Scientific Publications

04/05/2018
by   Andres Garcia, et al.
0

The emergence of knowledge graphs in the scholarly communication domain and recent advances in artificial intelligence and natural language processing bring us closer to a scenario where intelligent systems can assist scientists over a range of knowledge-intensive tasks. In this paper we present experimental results about the generation of word embeddings from scholarly publications for the intelligent processing of scientific texts extracted from SciGraph. We compare the performance of domain-specific embeddings with existing pre-trained vectors generated from very large and general purpose corpora. Our results suggest that there is a trade-off between corpus specificity and volume. Embeddings from domain-specific scientific corpora effectively capture the semantics of the domain. On the other hand, obtaining comparable results through general corpora can also be achieved, but only in the presence of very large corpora of well formed text. Furthermore, We also show that the degree of overlapping between knowledge areas is directly related to the performance of embeddings in domain evaluation tasks.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/11/2018

Domain Adapted Word Embeddings for Improved Sentiment Classification

Generic word embeddings are trained on large-scale generic corpora; Doma...
research
12/31/2022

Logic Mill – A Knowledge Navigation System

Logic Mill is a scalable and openly accessible software system that iden...
research
05/25/2018

Lifelong Domain Word Embedding via Meta-Learning

Learning high-quality domain word embeddings is important for achieving ...
research
09/21/2017

Learning Domain-Specific Word Embeddings from Sparse Cybersecurity Texts

Word embedding is a Natural Language Processing (NLP) technique that aut...
research
12/24/2021

Analyzing Scientific Publications using Domain-Specific Word Embedding and Topic Modelling

The scientific world is changing at a rapid pace, with new technology be...
research
07/01/2021

Leveraging Domain Agnostic and Specific Knowledge for Acronym Disambiguation

An obstacle to scientific document understanding is the extensive use of...
research
12/27/2022

TegFormer: Topic-to-Essay Generation with Good Topic Coverage and High Text Coherence

Creating an essay based on a few given topics is a challenging NLP task....

Please sign up or login with your details

Forgot password? Click here to reset