Seasonal-adjustment Based Feature Selection Method for Large-scale Search Engine Logs

08/22/2020
by   Thien Q. Tran, et al.
0

Search engine logs have a great potential in tracking and predicting outbreaks of infectious disease. More precisely, one can use the search volume of some search terms to predict the infection rate of an infectious disease in nearly real-time. However, conducting accurate and stable prediction of outbreaks using search engine logs is a challenging task due to the following two-way instability characteristics of the search logs. First, the search volume of a search term may change irregularly in the short-term, for example, due to environmental factors such as the amount of media or news. Second, the search volume may also change in the long-term due to the demographic change of the search engine. That is to say, if a model is trained with such search logs with ignoring such characteristic, the resulting prediction would contain serious mispredictions when these changes occur. In this work, we proposed a novel feature selection method to overcome this instability problem. In particular, we employ a seasonal-adjustment method that decomposes each time series into three components: seasonal, trend and irregular component and build prediction models for each component individually. We also carefully design a feature selection method to select proper search terms to predict each component. We conducted comprehensive experiments on ten different kinds of infectious diseases. The experimental results show that the proposed method outperforms all comparative methods in prediction accuracy for seven of ten diseases, in both now-casting and forecasting setting. Also, the proposed method is more successful in selecting search terms that are semantically related to target diseases.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
08/18/2020

Parallel Extraction of Long-term Trends and Short-term Fluctuation Framework for Multivariate Time Series Forecasting

Multivariate time series forecasting is widely used in various fields. R...
research
04/14/2021

Influenza Surveillance using Search Engine, SNS, On-line Shopping, Q A Service and Past Flu Patients

Influenza, an infectious disease, causes many deaths worldwide. Predicti...
research
06/11/2019

Medium-Term Load Forecasting Using Support Vector Regression, Feature Selection, and Symbiotic Organism Search Optimization

An accurate load forecasting has always been one of the main indispensab...
research
10/08/2017

Structural Feature Selection for Event Logs

We consider the problem of classifying business process instances based ...
research
05/24/2022

HiPAL: A Deep Framework for Physician Burnout Prediction Using Activity Logs in Electronic Health Records

Burnout is a significant public health concern affecting nearly half of ...
research
08/30/2019

Charge-Based Prison Term Prediction with Deep Gating Network

Judgment prediction for legal cases has attracted much research efforts ...

Please sign up or login with your details

Forgot password? Click here to reset