DeepAI AI Chat
Log In Sign Up

An Empirical Study of Sections in Classifying Disease Outbreak Reports

by   Son Doan, et al.

Identifying articles that relate to infectious diseases is a necessary step for any automatic bio-surveillance system that monitors news articles from the Internet. Unlike scientific articles which are available in a strongly structured form, news articles are usually loosely structured. In this chapter, we investigate the importance of each section and the effect of section weighting on performance of text classification. The experimental results show that (1) classification models using the headline and leading sentence achieve a high performance in terms of F-score compared to other parts of the article; (2) all section with bag-of-word representation (full text) achieves the highest recall; and (3) section weighting information can help to improve accuracy.


page 1

page 2

page 3

page 4


MN-DS: A Multilabeled News Dataset for News Articles Hierarchical Classification

This article presents a dataset of 10,917 news articles with hierarchica...

SaRoCo: Detecting Satire in a Novel Romanian Corpus of News Articles

In this work, we introduce a corpus for satire detection in Romanian new...

Using Word Embeddings to Analyze Protests News

The first two tasks of the CLEF 2019 ProtestNews events focused on disti...

Automated News Suggestions for Populating Wikipedia Entity Pages

Wikipedia entity pages are a valuable source of information for direct c...

KnowBias: A Novel AI Method to Detect Polarity in Online Content

We introduce KnowBias, a system for detecting the degree of political bi...

Suspicious News Detection Using Micro Blog Text

We present a new task, suspicious news detection using micro blog text. ...