RevDet: Robust and Memory Efficient Event Detection and Tracking in Large News Feeds

03/07/2021
by   Abdul Hameed Azeemi, et al.
0

With the ever-growing volume of online news feeds, event-based organization of news articles has many practical applications including better information navigation and the ability to view and analyze events as they develop. Automatically tracking the evolution of events in large news corpora still remains a challenging task, and the existing techniques for Event Detection and Tracking do not place a particular focus on tracking events in very large and constantly updating news feeds. Here, we propose a new method for robust and efficient event detection and tracking, which we call RevDet algorithm. RevDet adopts an iterative clustering approach for tracking events. Even though many events continue to develop for many days or even months, RevDet is able to detect and track those events while utilizing only a constant amount of space on main memory. We also devise a redundancy removal strategy which effectively eliminates duplicate news articles and substantially reduces the size of data. We construct a large, comprehensive new ground truth dataset specifically for event detection and tracking approaches by augmenting two existing datasets: w2e and GDELT. We implement RevDet algorithm and evaluate its performance on the ground truth event chains. We discover that our algorithm is able to accurately recover event chains in the ground-truth dataset. We also compare the memory efficiency of our algorithm with the standard single pass clustering approach, and demonstrate the appropriateness of our algorithm for event detection and tracking task in large news feeds.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/23/2021

Event Camera Based Real-Time Detection and Tracking of Indoor Ground Robots

This paper presents a real-time method to detect and track multiple mobi...
research
11/14/2019

Mining News Events from Comparable News Corpora: A Multi-Attribute Proximity Network Modeling Approach

We present ProxiModel, a novel event mining framework for extracting hig...
research
05/26/2021

Trade the Event: Corporate Events Detection for News-Based Event-Driven Trading

In this paper, we introduce an event-driven trading strategy that predic...
research
04/07/2018

Quootstrap: Scalable Unsupervised Extraction of Quotation-Speaker Pairs from Large News Corpora via Bootstrapping

We propose Quootstrap, a method for extracting quotations, as well as th...
research
12/12/2021

Topic Detection and Tracking with Time-Aware Document Embeddings

The time at which a message is communicated is a vital piece of metadata...
research
10/12/2021

Topic-time Heatmaps for Human-in-the-loop Topic Detection and Tracking

The essential task of Topic Detection and Tracking (TDT) is to organize ...
research
09/15/2019

Query-Focused Scenario Construction

The news coverage of events often contains not one but multiple incompat...

Please sign up or login with your details

Forgot password? Click here to reset