Skip to content
Course contents
Testing an edge honestly

Torture the data and it confesses

This is the deepest honesty chapter in the catalogue. If you test enough strategies, the best one will look excellent purely by chance, so a great backtest found after many tries is probably luck, not skill. This chapter demonstrates the effect directly in code and names the trap that ruins most quant research.

10 min readChapter 16 of 26
What you will learn
  • Explain the multiple-testing problem and selection bias
  • Demonstrate in code that searching many random strategies yields an impressive false winner
  • Track the number of trials as a first line of defence

This is the most important chapter in the course, and the one most quants learn too late. If you test enough strategies, one of them will look brilliant purely by chance, and you will not be able to tell it from a real edge by looking at its backtest. This is the multiple-testing problem, and it is the deepest reason quant research fools people. The old joke is exact: torture the data long enough and it will confess to anything.

Searching manufactures winners

To see it, we do something deliberately hopeless: generate strategies with no edge at all, pure noise with a true average return of zero, then search many of them and keep the best.

ExampleSearching many no-edge strategies manufactures an impressive bestch16/multiple_testing.py
# The multiple-testing problem. We generate strategies with NO real edge (pure
# noise, true mean zero), search many of them, and keep the best. The more we try,
# the better the best one looks, purely by luck. This is why a great backtest found
# after many tries is probably nothing.
import numpy as np
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt

rng = np.random.default_rng(5)
T = 120                     # 10 years of monthly returns: a generous backtest


def best_sharpe_of(n_trials):
    r = rng.normal(0, 0.04, (n_trials, T))              # all no-edge
    s = r.mean(axis=1) / r.std(axis=1, ddof=1) * np.sqrt(12)
    return s.max()


print("Strategies tried  ->  best Sharpe found (every one has NO real edge):")
for n in [1, 10, 100, 1000, 5000]:
    print(f"  {n:>5} tried  ->  best Sharpe {best_sharpe_of(n):.2f}")

# The distribution of the best-of-1000 Sharpe over many independent searches.
best_of_1000 = np.array([best_sharpe_of(1000) for _ in range(500)])
print(f"\nSearching 1000 no-edge strategies typically yields a best Sharpe near "
      f"{best_of_1000.mean():.2f},")
print("a level practitioners would call a genuine edge, produced here by pure luck.")

plt.figure(figsize=(7, 4))
plt.hist(best_of_1000, bins=30)
plt.axvline(0, color="grey", linewidth=0.8)
plt.title("The best of 1000 no-edge strategies looks like an edge, by luck")
plt.xlabel("Best Sharpe found among 1000 random strategies")
plt.ylabel("Frequency (over 500 searches)")
plt.tight_layout()
plt.savefig("multiple_testing.png", dpi=110)
print("Saved multiple_testing.png")
Output
Strategies tried  ->  best Sharpe found (every one has NO real edge):
      1 tried  ->  best Sharpe -0.71
     10 tried  ->  best Sharpe 0.54
    100 tried  ->  best Sharpe 0.68
   1000 tried  ->  best Sharpe 0.90
   5000 tried  ->  best Sharpe 1.22

Searching 1000 no-edge strategies typically yields a best Sharpe near 1.05,
a level practitioners would call a genuine edge, produced here by pure luck.
Saved multiple_testing.png
Searching many no-edge strategies manufactures an impressive best generated from the code above

Follow the numbers, remembering that not one of these strategies has any real edge. A single random strategy has a Sharpe scattered around zero, here minus 0.71. But search ten and the best is 0.54; search a hundred and it is 0.68; search a thousand and it is 0.90; search five thousand and it is 1.22. The distribution of the best-of-a-thousand clusters around 1.05, and the chart shows it never coming anywhere near the true value of zero. A Sharpe near 1 is a level a practitioner would call a genuine edge, and here it has been conjured entirely from noise, just by searching. The best backtest from a large search is not evidence of skill. It is the expected result of searching.

Why it is so dangerous

What makes this lethal is that it is invisible. The winning strategy's backtest looks identical to a real edge's: the same smooth equity curve, the same impressive Sharpe. Nothing on the surface reveals that it was the luckiest of a thousand tries. And multiple testing is everywhere in practice, usually unrecognised. Every optimiser that sweeps parameters, every "I tried a few variations and kept the best one", every grid search over lookbacks and thresholds, is multiple testing, and the harder you search, the more inflated the winner. This is why a beautiful backtest is such weak evidence on its own, and why the number of things you tried matters as much as the result you found.

The first defence: count your trials

The first line of defence is simple and almost never practised: be honest about how many strategies you tried, because you cannot correct for a search you refuse to count. Keep a record of every variation tested, every parameter swept, every idea run. That count is not bookkeeping; it is the input to the correction the deflated Sharpe chapter applies. The next chapter adds the validation defence, testing on unseen data, and the chapter after turns the count of trials into an actual haircut on your Sharpe. For now, hold the humbling fact: with enough tries, an excellent backtest is guaranteed, whether or not any edge exists.

What to carry forward

The multiple-testing problem is the deepest trap in quant research: search enough strategies and the best one looks brilliant by pure chance, as you saw a thousand no-edge strategies reliably yield a best Sharpe near 1. Its backtest is indistinguishable from a real edge, and every optimiser and parameter sweep is a form of it, so a beautiful backtest is weak evidence and the number of trials matters as much as the result. The first defence is to count your trials honestly. The next chapter adds the second and stronger defence: judging a chosen strategy only on data it has never seen.