Skip to content

Why research runs before the model

A model fitted on results cannot know a goalkeeper was ruled out this morning. Research is how information that has not reached the data gets into the number.

BAI8 ResearchResearch desk
2 min read

The fitted half of this product is blind by construction. It knows what teams with these rates have done against teams with those rates, over years of results, and that is a strong base. It does not know that the first-choice goalkeeper failed a fitness test four hours ago, because no result has been recorded that contains that fact.

What the passes actually look for

Thirteen agents read the fixture from thirteen angles: the official team news, the pre-match press conference, local reporting in the local language, the injury history of anybody doubtful, the fixture congestion either side, travel distance, and what the market has already done in the last day.

That last one matters more than it sounds. If a price has moved sharply and the research finds nothing to explain it, that is itself a finding: it usually means somebody knows something we have not found yet, and the honest response is to widen the uncertainty rather than to insist.

Sourced or discarded

Every claim comes back with a source attached, and one without a source is thrown away rather than kept at a lower weight. This is stricter than it needs to be and it is deliberate. The moment unsourced claims are allowed in at low weight, the number becomes impossible to audit: when it turns out wrong, there is no way to point at the step that made it wrong.

Where it joins the model

The two halves are composed rather than averaged, weighted by how much the research actually found. A fixture where nothing was learned stays close to the model, which is the correct answer. The full order is set out in how BAI8 prices a fixture.

What this does not tell you

Research is not a superpower. It finds what is published, and what is published about a mid-table fixture in a minor league is thin to the point of useless. It also has a correlated failure mode the model does not: when the press is wrong about a team, it is wrong about that team everywhere, and every fixture involving them moves the same wrong way.

Common questions

Why does a betting model need research at all?

Because a model fitted on historical results only knows what has already been recorded. Team news, travel and late fitness decisions have not been recorded anywhere the model can see.

How is a research claim verified?

Every claim carries the source it came from. A claim without a source is discarded rather than downweighted, because an unsourced claim cannot be checked later when the number turns out to be wrong.

Can research make a number worse?

Yes. Three outlets repeating one rumour looks like three sources, and weighting it as three is a real failure mode rather than a hypothetical one.

Sources

  1. Bayesian inferenceen.wikipedia.org
  • research
  • model
  • team news
  • sources

Part of

Read next