A Calibrated Data-Driven Approach for Small Area Estimation using Big Data

06/06/2023
by   Siu-Ming Tam, et al.
0

Where the response variable in a big data set is consistent with the variable of interest for small area estimation, the big data by itself can provide the estimates for small areas. These estimates are often subject to the coverage and measurement error bias inherited from the big data. However, if a probability survey of the same variable of interest is available, the survey data can be used as a training data set to develop an algorithm to impute for the data missed by the big data and adjust for measurement errors. In this paper, we outline a methodology for such imputations based on an kNN algorithm calibrated to an asymptotically design-unbiased estimate of the national total and illustrate the use of a training data set to estimate the imputation bias and the fixed - asymptotic bootstrap to estimate the variance of the small area hybrid estimator. We illustrate the methodology of this paper using a public use data set and use it to compare the accuracy and precision of our hybrid estimator with the Fay-Harriot (FH) estimator. Finally, we also examine numerically the accuracy and precision of the FH estimator when the auxiliary variables used in the linking models are subject to under-coverage errors

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/26/2020

Data Integration by combining big data and survey sample data for finite population inference

The statistical challenges in using big data for making valid statistica...
research
06/28/2023

Integrating Big Data and Survey Data for Efficient Estimation of the Median

An ever-increasing deluge of big data is becoming available to national ...
research
07/22/2023

Survey Design and Estimating Equations when Combining Big Data with Probability Samples

The use of big data in official statistics and the applied sciences is a...
research
03/06/2021

Transformed Fay-Herriot Model with Measurement Error in Covariates

Statistical agencies are often asked to produce small area estimates (SA...
research
06/15/2019

Proxy expenditure weights for Consumer Price Index: Audit sampling inference for big data statistics

Purchase data from retail chains provide proxy measures of private house...
research
08/14/2022

Sharp Frequency Bounds for Sample-Based Queries

A data sketch algorithm scans a big data set, collecting a small amount ...
research
08/05/2018

Mining CFD Rules on Big Data

Current conditional functional dependencies (CFDs) discovery algorithms ...

Please sign up or login with your details

Forgot password? Click here to reset