Skip to content

A bad week: what the results can and cannot show

A worked audit of nine assessments separates ordinary outcome variance from a process failure, without presenting one bad week as proof.

BAI8 Research21 Aug 2026, updated 25 Sept 20262 min read

On this page
  1. Eight of the nine were fine
  2. One was not fine
  3. Why the distinction is the whole point
  4. What this does not tell you

Take a worked sample of nine assessments with seven losses. Those figures are illustrative, not BAI8's live performance record. Without the probability assigned to each event, it is not possible to state honestly how surprising the sequence is. It is enough to show why results and process must be audited separately.

Eight of the nine were fine

By fine we mean the reasoning holds up on re-reading, the sources are real, and the number would be the same if it were reproduced knowing only what was available at the time. Seven of those eight lose. That is consistent with what variance does and there is nothing specific to fix from the outcomes alone.

Suppose six of the eight are taken at a better price than the market closes at. That adds a price comparison to the review, but it still does not make nine observations enough to validate a method.

One was not fine

In the example, a team sheet is published two hours before kick-off and then amended. The research pass reads the first version. Nothing downstream checks the timestamp, so an assessment is built on a lineup that has already been superseded. That is a demonstrable input error rather than an inference from the result.

That process failure has a testable response: store the publication time with each claim and refetch it when the same origin posts a newer update. The fix is not proved by describing it; it needs controlled tests and subsequent audits.

Why the distinction is the whole point

If bad weeks and broken processes are counted the same way, then every bad week produces a change and the method drifts toward whatever the last few results suggested. A proper scoring rule evaluates the probabilities across a series rather than rewarding a reactive change after every loss.

Separating them requires a complete, timestamped record that includes the weeks a publisher would rather omit.

What this does not tell you

Whether the timestamp response is sufficient. It addresses the example we defined. It does not estimate how often similar failures occurred before, or how often they happened to settle in the publisher's favour.

Common questions

Is this a report of BAI8's live performance?

No. It is a worked example of how to review a bad week without confusing outcomes with process quality.

Do seven losses from nine prove the assessments are broken?

No. The probability of that sequence depends on the probability assigned to each event, and nine observations are too few to diagnose a method by win rate alone.

What process error does the example use?

A stale team sheet. The example assumes a research pass read a lineup published before a late change and nothing downstream noticed the timestamp.

Sources

  1. NIST exact confidence intervals for a binomial proportionitl.nist.gov
  2. Verification of Forecasts Expressed in Terms of Probabilityjournals.ametsoc.org
  3. Strictly Proper Scoring Rules, Prediction, and Estimationtandfonline.com

Read next