Estimation of finite population proportions for small areas: a statistical data integration approach

05/21/2023
by   Aditi Sen, et al.
0

Empirical best prediction (EBP) is a well-known method for producing reliable proportion estimates when the primary data source provides only small or no sample from finite populations. There are at least two potential challenges encountered in implementing the existing EBP methodology. First, one must accurately link the sample to the finite population frame. This may be a difficult or even impossible task because of absence of identifiers that can be used to link sample and the frame. Secondly, the finite population frame typically contains limited auxiliary variables, which may not be adequate for building a reasonable working predictive model. We propose a data linkage approach in which we replace the finite population frame by a big sample that does not have the outcome binary variable of interest, but has a large set of auxiliary variables. Our proposed method calls for fitting the assumed model using data from the smaller sample, imputing the outcome variable for all the units of the big sample, and then finally using these imputed values to obtain standard weighted proportion using the big sample. We develop a new adjusted maximum likelihood method to avoid estimates of model variance on the boundary encountered in the commonly used in maximum likelihood estimation method. We propose an estimator of mean squared prediction error (MSPE) using a parametric bootstrap method and address computational issues by developing efficient EM algorithm. We illustrate the proposed methodology in the context of election projection for small areas.

READ FULL TEXT

page 25

page 29

research
10/10/2022

Hierarchical Bayes estimation of small area proportions using statistical linkage of disparate data sources

We propose a Bayesian approach to estimate finite population proportions...
research
03/24/2019

Cost Issue in Estimation of Proportion in a Finite Population Divided Among Two Strata

The problem of estimation of the proportion of units with a given attrib...
research
03/16/2019

An example of application of optimal sample allocation in a finite population

The problem of estimating a proportion of objects with particular attrib...
research
03/24/2021

Statistical Integration of Heterogeneous Data with PO2PLS

The availability of multi-omics data has revolutionized the life science...
research
12/10/2019

Variable selection for transportability

Transportability provides a principled framework to address the problem ...
research
10/15/2021

Estimating individual admixture from finite reference databases

The concept of individual admixture (IA) assumes that the genome of indi...
research
05/11/2021

Estimation of mask effectiveness perception for small domains using multiple data sources

All pandemics are local; so learning about the impacts of pandemics on p...

Please sign up or login with your details

Forgot password? Click here to reset