The month in full: how to audit 41 assessments
A worked audit of 41 settled assessments shows why outcomes, closing prices and process errors must be reported separately before drawing conclusions.
On this page
Consider a worked monthly audit of forty-one assessments. Twenty-seven win, fourteen lose, and the month finishes ahead. Those figures are illustrative, not a claim about BAI8's live record. The useful question is what a complete audit could establish and what it could not.
The number that carries information
In this example, thirty-two of the forty-one are taken at a price better than the one the market settles on. That comparison is worth recording, for the reasons set out in why the closing line is the only honest scoreboard: it is measured on every bet immediately and it does not care whether the ball went in.
Nine close the other way. Six of those nine are fixtures where the research finds nothing and the assessed number stays close to the market price it started from. That pattern is a reason to investigate, not evidence by itself that the market later found something the research missed.
Two where the reasoning was wrong
Distinguishing process errors from ordinary losses matters because only the former identify a specific part of the method that needs correction.
Suppose the first is a totals market where the research reads a scoring rate from a run of fixtures against three promoted sides. That is a sampling problem the review should catch. The correction would need an out-of-sample test before it could be treated as an improvement.
Suppose the second is a research failure of the kind described in why no number is published before the research: three outlets carried the same injury claim, it was counted as three sources, and it originated from one. That calls for source-level deduplication, followed by a test that the change catches duplicates without discarding independent confirmation.
What a good month is worth
Very little on its own. Exact binomial confidence intervals are wide at small sample sizes. Without the probability assigned to each assessment, even that calculation is incomplete. Reading one profitable month as vindication would be exactly the error this piece exists to avoid.
What this does not tell you
Whether either proposed correction makes the assessments better. Each responds to the example that produced it, which is a different and much weaker claim. The honest test is whether the change improves probability scores out of sample.
Common questions
Is this BAI8's live performance record?
No. It is a worked example showing how a complete monthly audit should separate results, prices and process errors.
Is a profitable month evidence that a method works?
No. A month is a small sample, so its win rate can sit far from the underlying rate by chance alone.
What should happen when the reasoning was wrong rather than unlucky?
Record it separately. A loss with sound reasoning and a loss caused by a process mistake are different events and should not be diagnosed in the same way.