Estimating the number of unseen species: A bird in the hand is worth n in the bush

by   Alon Orlitsky, et al.

Estimating the number of unseen species is an important problem in many scientific endeavors. Its most popular formulation, introduced by Fisher, uses n samples to predict the number U of hitherto unseen species that would be observed if t· n new samples were collected. Of considerable interest is the largest ratio t between the number of new and existing samples for which U can be accurately predicted. In seminal works, Good and Toulmin constructed an intriguing estimator that predicts U for all t< 1, thereby showing that the number of species can be estimated for a population twice as large as that observed. Subsequently Efron and Thisted obtained a modified estimator that empirically predicts U even for some t>1, but without provable guarantees. We derive a class of estimators that provably predict U not just for constant t>1, but all the way up to t proportional to n. This shows that the number of species can be estimated for a population n times larger than that observed, a factor that grows arbitrarily large as n increases. We also show that this range is the best possible and that the estimators' mean-square error is optimal up to constants for any t. Our approach yields the first provable guarantee for the Efron-Thisted estimator and, in addition, a variant which achieves stronger theoretical and experimental performance than existing methodologies on a variety of synthetic and real datasets. The estimators we derive are simple linear estimators that are computable in time proportional to n. The performance guarantees hold uniformly for all distributions, and apply to all four standard sampling models commonly used across various scientific disciplines: multinomial, Poisson, hypergeometric, and Bernoulli product.


page 1

page 2

page 3

page 4


Near-optimal estimation of the unseen under regularly varying tail populations

Given n samples from a population of individuals belonging to different ...

A Good-Turing estimator for feature allocation models

Feature allocation models generalize species sampling models by allowing...

Estimating the unseen from multiple populations

Given samples from a distribution, how many new elements should we expec...

Convergence of Chao Unseen Species Estimator

Support size estimation and the related problem of unseen species estima...

Richness estimation with species identity error

Richness estimation of an interesting area is always a challenge statist...

Reducing Seed Bias in Respondent-Driven Sampling by Estimating Block Transition Probabilities

Respondent-driven sampling (RDS) is a popular approach to study marginal...

Support Estimation via Regularized and Weighted Chebyshev Approximations

We introduce a new framework for estimating the support size of an unkno...

Please sign up or login with your details

Forgot password? Click here to reset