Metrics of calibration for probabilistic predictions

05/19/2022
by   Imanol Arrieta Ibarra, et al.
0

Predictions are often probabilities; e.g., a prediction could be for precipitation tomorrow, but with only a 30 predictions together with the actual outcomes, "reliability diagrams" help detect and diagnose statistically significant discrepancies – so-called "miscalibration" – between the predictions and the outcomes. The canonical reliability diagrams histogram the observed and expected values of the predictions; replacing the hard histogram binning with soft kernel density estimation is another common practice. But, which widths of bins or kernels are best? Plots of the cumulative differences between the observed and expected values largely avoid this question, by displaying miscalibration directly as the slopes of secant lines for the graphs. Slope is easy to perceive with quantitative precision, even when the constant offsets of the secant lines are irrelevant; there is no need to bin or perform kernel density estimation. The existing standard metrics of miscalibration each summarize a reliability diagram as a single scalar statistic. The cumulative plots naturally lead to scalar metrics for the deviation of the graph of cumulative differences away from zero; good calibration corresponds to a horizontal, flat graph which deviates little from zero. The cumulative approach is currently unconventional, yet offers many favorable statistical properties, guaranteed via mathematical theory backed by rigorous proofs and illustrative numerical examples. In particular, metrics based on binning or kernel density estimation unavoidably must trade-off statistical confidence for the ability to resolve variations as a function of the predicted probability or vice versa. Widening the bins or kernels averages away random noise while giving up some resolving power. Narrowing the bins or kernels enhances resolving power while not averaging away as much noise.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
06/03/2020

Plots of the cumulative differences between observed and expected values of ordered Bernoulli variates

Many predictions are probabilistic in nature; for example, a prediction ...
research
08/04/2020

Plotting the cumulative deviation of a subgroup from the full population as a function of score

Assessing whether a subgroup of a full population is getting treated equ...
research
08/05/2021

Cumulative differences between subpopulations

Comparing the differences in outcomes (that is, in "dependent variables"...
research
01/31/2022

Calibration of P-values for calibration and for deviation of a subpopulation from the full population

The author's recent research papers, "Cumulative deviation of a subpopul...
research
09/21/2023

Smooth ECE: Principled Reliability Diagrams via Kernel Smoothing

Calibration measures and reliability diagrams are two fundamental tools ...
research
05/07/2020

Fast multivariate empirical cumulative distribution function with connection to kernel density estimation

This paper revisits the problem of computing empirical cumulative distri...
research
03/07/2018

Nonparametric Estimation of Probability Density Functions of Random Persistence Diagrams

We introduce a nonparametric way to estimate the global probability dens...

Please sign up or login with your details

Forgot password? Click here to reset