Data-snooping bias is the error of testing many strategy variants against the same data until one of them looks good, then reporting that winner as though it were a genuine finding. The winner's headline figures are real; what has been lost is the information needed to judge them, because the result was not one test but the best of many, and the best of many looks impressive even when nothing at all is there. Data-snooping bias is a property of the search that produced a result, not of the result itself — which is why it is invisible in the report and can only be assessed by knowing how many candidates were tried.
Also seen as: Data mining bias, multiple-testing bias, p-hacking
Why does testing many ideas produce a false winner?
Testing many ideas produces a false winner because each individual test carries its own chance of looking good purely by luck, and those chances accumulate across the search. If every test has a probability of appearing significant when there is no real effect, then the chance that at least one of independent tests does so is:
where: is the false-positive rate of a single test, is the number of independent tests run, and the result is the familywise error rate — the chance the search produces at least one apparently significant result from pure noise.
The consequence is that the meaning of a threshold depends on how many times it was applied. A result that would be notable as the outcome of one pre-registered test is unremarkable as the best of two hundred, and the same numbers on the page therefore support very different conclusions depending on a fact the page does not contain. In quantitative finance the standard corrections either raise the bar for significance as the number of trials grows (as the Bonferroni correction does) or test the surviving candidate on data it has never seen.
Worked example: 100 ideas, one winner
Take — a one-in-twenty chance that any single test looks good on noise alone.
| Ideas tested () | Chance at least one looks good by luck |
|---|---|
| 1 | 5% |
| 14 | 51% |
| 100 | 99.4% |
Run 100 unrelated strategy variants and it is effectively certain — 99.4% — that at least one clears the bar for no reason at all. Report only that variant and you have a page of strong figures produced by a coin-flipping exercise. The arithmetic also shows how quickly this arrives: by the fourteenth variant the odds are already worse than even. Note what is not recorded anywhere in the winner's own results: the other 99 runs. A reader given only the winner cannot compute the row of this table that applies.
How is data-snooping bias different from overfitting?
Data-snooping bias and overfitting are the same problem approached from opposite directions. Overfitting is about the strategy — it has enough degrees of freedom to absorb noise, so it describes one sample too precisely. Data-snooping is about the search — many candidates were evaluated and only the survivor is shown, so the survivor's quality partly measures the size of the search. A single hand-tuned twelve-parameter strategy is overfitted with no search at all; a one-parameter strategy picked as the best of 500 tests is data-snooped without being complex.
It is also distinct from its reporting-stage neighbour: cherry-picking bias is quoting the favorable part of one result, while data-snooping is quoting the favorable result out of many. And it differs from selection bias, which is about an unrepresentative sample of instruments or periods rather than a repeated search.
What does Fincanva do about data-snooping bias?
Fincanva does not prevent data-snooping, and no backtesting tool can: the app lets you run as many variants as you like, and a completed backtest carries no record of how many other variants you tried before it. Keeping track of the size of your own search is therefore left to you, and it is the only input that makes a winner interpretable. That count also has to survive confirmation bias, which is the pull to remember the runs that agreed with the idea and quietly discount the rest.
Two product behaviors bear on it. A backtest reports the full-period result with its drawdowns, not only the favorable stretch, so a lucky winner is at least shown in full. And the start-date sensitivity view re-runs one strategy across many entry dates and reports the spread of outcomes, which tends to expose a result that survived only one particular window. See the nine biases Fincanva helps you avoid for the product behavior in context.
Backtests show what would have happened — not what will. Fincanva provides no financial advice — see Is this financial advice?.
Where this term is used
Generated · 1 pagesThe pages that reference this term — so a term page is somewhere you pass through, not somewhere you land and stop.
Also referenced by 3 terms
Fincanva provides no financial advice. Backtests show what would have happened — not what will.
GLOSSARY · 193 TERMS