Skip to content

What a calibration curve shows

A model that says 70% should be right about seven times in ten. The curve that checks this is the plainest test of a forecast there is, and it is easy to read.

BAI8 Research16 Jul 2026, updated 25 Sept 20262 min read

On this page
  1. Reading the plot
  2. Why calibration alone is not enough
  3. Where this sits in the pipeline
  4. What this does not tell you

Take every fixture where a model said 70%, and count how many of them happened. If the answer is about seven in ten, that number means what it says. If the answer is five in ten, the forecast is miscalibrated in that region. The curve alone does not establish whether the model ranked the sides in the right order; it shows that the stated probability did not match the observed frequency.

Reading the plot

Predicted probability runs along the bottom, observed frequency up the side, and a perfectly calibrated model traces the diagonal. Points above the line mean the model is underconfident: things it called 40% happen half the time. Points below mean it is overconfident. The distance from the line must be read alongside the number of forecasts in that region and their sampling uncertainty.

Why calibration alone is not enough

A model that answers "the home side wins 46% of the time" for every fixture in the league can be perfectly calibrated if home sides win 46% of that sample. It still has no resolution across matches because it never says anything about a specific fixture. The second property a forecast needs is sharpness: the willingness to move away from the base rate when there is a reason to.

Calibration without sharpness is a weather forecaster who always says "average". Sharpness without calibration is one who always says "certain". Both are useless in different directions, and both look fine on the metric the other one fails.

Where this sits in the pipeline

Calibration is checked on the assessed number, because that is the only number this product publishes. Each assessment keeps its market starting point and the research adjustment on record separately. That matters: if the published number is miscalibrated, it is possible to ask whether the starting price or the research introduced it, and the whole point of the pipeline is that every step can be blamed separately.

What this does not tell you

A calibration curve needs enough observations in each probability region. With sparse bins, observed frequencies have wide sampling uncertainty, and arbitrary bin choices can create apparent wobbles. It also says nothing about whether the model beats the market, which is a different and harder question.

Common questions

What is a calibration curve?

A plot of predicted probability against observed frequency. Group every 70% forecast together, count how often those events happened, and plot the pair. A perfectly calibrated model sits on the diagonal.

Does good calibration mean a model is useful?

Not by itself. A model that predicts the observed base rate for every match can be calibrated in aggregate while having no resolution across fixtures.

What is the difference between calibration and accuracy?

Calibration asks whether stated probabilities match observed frequencies. Classification accuracy asks whether thresholded labels were correct, while sharpness or resolution asks whether forecasts distinguish cases. They are separate properties.

Sources

  1. Dimitriadis, Gneiting and Jordan, Stable reliability diagrams for probabilistic classifiersdoi.org
  2. Gneiting, Balabdaoui and Raftery, Probabilistic Forecasts, Calibration and Sharpnessdoi.org

Read next