ReadNet: A Hierarchical Transformer Framework for Web Article Readability Analysis

03/06/2021
by   Changping Meng, et al.
0

Analyzing the readability of articles has been an important sociolinguistic task. Addressing this task is necessary to the automatic recommendation of appropriate articles to readers with different comprehension abilities, and it further benefits education systems, web information systems, and digital libraries. Current methods for assessing readability employ empirical measures or statistical learning techniques that are limited by their ability to characterize complex patterns such as article structures and semantic meanings of sentences. In this paper, we propose a new and comprehensive framework which uses a hierarchical self-attention model to analyze document readability. In this model, measurements of sentence-level difficulty are captured along with the semantic meanings of each sentence. Additionally, the sentence-level features are incorporated to characterize the overall readability of an article with consideration of article structures. We evaluate our proposed approach on three widely-used benchmark datasets against several strong baseline approaches. Experimental results show that our proposed method achieves the state-of-the-art performance on estimating the readability for various web articles and literature.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/28/2020

HIN: Hierarchical Inference Network for Document-Level Relation Extraction

Document-level RE requires reading, inferring and aggregating over multi...
research
04/18/2021

On the Use of Context for Predicting Citation Worthiness of Sentences in Scholarly Articles

In this paper, we study the importance of context in predicting the cita...
research
09/12/2022

Large-scale Evaluation of Transformer-based Article Encoders on the Task of Citation Recommendation

Recently introduced transformer-based article encoders (TAEs) designed t...
research
11/08/2016

Sentence Ordering using Recurrent Neural Networks

Modeling the structure of coherent texts is a task of great importance i...
research
11/08/2019

Question Generation from Paragraphs: A Tale of Two Hierarchical Models

Automatic question generation from paragraphs is an important and challe...
research
11/16/2021

WikiContradiction: Detecting Self-Contradiction Articles on Wikipedia

While Wikipedia has been utilized for fact-checking and claim verificati...
research
03/23/2021

A General Framework for Learning Prosodic-Enhanced Representation of Rap Lyrics

Learning and analyzing rap lyrics is a significant basis for many web ap...

Please sign up or login with your details

Forgot password? Click here to reset