Skip to content

What a calibration curve shows

A model that says 70% should be right about seven times in ten. The curve that checks this is the plainest test of a forecast there is, and it is easy to read.

BAI8 ResearchResearch desk
2 min read

Take every fixture where a model said 70%, and count how many of them happened. If the answer is about seven in ten, that number means what it says. If the answer is five in ten, the model is not wrong about which side is better, it is wrong about how sure it is, and that is a separate and more correctable fault.

Reading the plot

Predicted probability runs along the bottom, observed frequency up the side, and a perfectly calibrated model traces the diagonal. Points above the line mean the model is underconfident: things it called 40% happen half the time. Points below mean it is overconfident, which is the more common failure and the more expensive one, because overconfidence at long prices is where money disappears fastest.

Why calibration alone is not enough

A model that answers "the home side wins 46% of the time" for every fixture in the league will be beautifully calibrated. It will also be worthless, because it never says anything about a specific match. The second property a forecast needs is sharpness: the willingness to move away from the base rate when there is a reason to.

Calibration without sharpness is a weather forecaster who always says "average". Sharpness without calibration is one who always says "certain". Both are useless in different directions, and both look fine on the metric the other one fails.

Where this sits in the pipeline

Calibration is checked on the fitted half, before research touches it. That ordering matters: if the composed number is miscalibrated it is much harder to say whether the model or the research introduced it, and the whole point of the pipeline is that every step can be blamed separately.

What this does not tell you

A calibration curve needs volume. Fitted on two hundred matches it is mostly noise, and reading a wobble in it as a finding is how people talk themselves into changes that make a model worse. It also says nothing about whether the model beats the market, which is a different and harder question.

Common questions

What is a calibration curve?

A plot of predicted probability against observed frequency. Group every 70% forecast together, count how often those events happened, and plot the pair. A perfectly calibrated model sits on the diagonal.

Does good calibration mean a model is useful?

Not by itself. A model that predicts the base rate for every match is perfectly calibrated and completely useless, because it never distinguishes one fixture from another.

What is the difference between calibration and accuracy?

Calibration asks whether the numbers mean what they say. Sharpness asks whether they say anything specific. A useful forecast needs both.

Sources

  1. Calibration in classificationen.wikipedia.org
  2. Brier scoreen.wikipedia.org
  • calibration
  • model
  • evaluation
  • probability

Part of

Read next