Skip to content
Course contents
Testing an edge honestly

The honest Sharpe ratio

There is a way to put a number on the multiple-testing problem: the deflated Sharpe ratio, which lowers a strategy's apparent Sharpe to account for how many strategies you tried before finding it. This chapter explains the idea in plain terms, computes it in code, and uses it to turn an impressive backtest into an honest expectation.

10 min readChapter 18 of 26
What you will learn
  • Explain why the number of trials inflates the best Sharpe ratio found
  • Compute a deflated Sharpe ratio that corrects for the number of trials and sample length
  • Use it to judge whether a discovered strategy is likely real

Out-of-sample testing catches a false edge, but it does not put a number on the multiple-testing problem. There is a tool that does: the deflated Sharpe ratio, developed by Bailey and Lopez de Prado, which lowers a strategy's apparent Sharpe to account for how many strategies you tried before finding it. It turns "I searched a thousand and this was the best" into an honest expectation, and it is the most grown-up number in quant research.

The haircut

Try enough strategies and one will look great by luck; the deflated Sharpe ratio lowers a strategy's apparent Sharpe for how many you tried. Illustrative.
Try enough strategies and one will look great by luck; the deflated Sharpe ratio lowers a strategy's apparent Sharpe for how many you tried. Illustrative.

The idea rests directly on the multiple-testing chapter. The more strategies you try, the higher the best Sharpe you would expect to find from pure luck, so a discovered Sharpe has to clear that luck bar before it counts for anything. That expected best from luck is the haircut, and it grows with the number of trials.

ExampleThe deflation haircut: the expected best Sharpe from N no-edge trialsch18/deflated_sharpe.py
# The deflated Sharpe ratio idea: if you tried N strategies, the bar for a Sharpe
# that is "probably real" rises, because the best of N no-edge strategies already
# has a high Sharpe by luck. We estimate that expected best (the haircut) two ways,
# by a formula and by simulation, then use it to judge an observed Sharpe.
import numpy as np
from scipy import stats

rng = np.random.default_rng(21)
T = 120                                 # months of data (matches the ch16 search)
se_ann = np.sqrt(12.0 / T)              # approximate SE of an annualised Sharpe near zero
gamma = 0.5772156649                    # Euler-Mascheroni constant


def expected_max_formula(n_trials):
    # Expected maximum of n_trials standard normals (Gumbel approximation), scaled
    # by the standard error of the Sharpe. This is the deflation "haircut".
    e_max_z = ((1 - gamma) * stats.norm.ppf(1 - 1.0 / n_trials)
               + gamma * stats.norm.ppf(1 - 1.0 / (n_trials * np.e)))
    return se_ann * e_max_z


def expected_max_sim(n_trials, reps=3000):
    out = []
    for _ in range(reps):
        r = rng.normal(0, 0.04, (n_trials, T))
        s = r.mean(axis=1) / r.std(axis=1, ddof=1) * np.sqrt(12)
        out.append(s.max())
    return np.mean(out)


print(f"With {T} months, the SE of an annualised Sharpe is about {se_ann:.2f}.")
print("(With a single strategy there is no haircut: your Sharpe stands as measured.)")
print("Trials  ->  expected BEST Sharpe from no-edge strategies (the haircut):")
for n in [10, 100, 1000]:
    print(f"  {n:>4}:  formula {expected_max_formula(n):.2f},  simulation {expected_max_sim(n):.2f}")

# Judge an observed strategy against the haircut for the number you tried.
observed, n_tried = 1.0, 1000
haircut = expected_max_formula(n_tried)
print(f"\nYou found a Sharpe of {observed:.1f} after trying {n_tried} strategies.")
print(f"The expected best from {n_tried} no-edge tries is {haircut:.2f} (the haircut).")
if observed > haircut:
    print("It clears the haircut: some evidence of a real edge, not just search luck.")
else:
    print("It does NOT clear the haircut: a Sharpe of 1.0 is within what luck alone")
    print("produces from 1000 tries, so the deflated Sharpe is effectively zero.")
Output
With 120 months, the SE of an annualised Sharpe is about 0.32.
(With a single strategy there is no haircut: your Sharpe stands as measured.)
Trials  ->  expected BEST Sharpe from no-edge strategies (the haircut):
    10:  formula 0.50,  simulation 0.49
   100:  formula 0.80,  simulation 0.81
  1000:  formula 1.03,  simulation 1.05

You found a Sharpe of 1.0 after trying 1000 strategies.
The expected best from 1000 no-edge tries is 1.03 (the haircut).
It does NOT clear the haircut: a Sharpe of 1.0 is within what luck alone
produces from 1000 tries, so the deflated Sharpe is effectively zero.

With ten years of monthly data, the standard error of a Sharpe estimate is about 0.32, which sets the scale of the luck. From that, the expected best Sharpe among ten no-edge strategies is 0.50, among a hundred it is 0.80, and among a thousand it is 1.03. The formula and a direct simulation agree closely, which is a good sign the formula is trustworthy. And notice that this haircut of 1.03 for a thousand trials is exactly the level the multiple-testing chapter measured when it searched a thousand no-edge strategies and found a best Sharpe near 1.05. The deflated Sharpe is simply naming, with a formula, the luck that chapter demonstrated.

Using it

Now judge a real result. Suppose you searched a thousand strategies and your best had a Sharpe of 1.0. It looks like a genuine edge. But the expected best from a thousand no-edge tries is 1.03, so your 1.0 does not even clear the haircut: it is squarely within what luck alone produces from that many attempts, and its deflated Sharpe is effectively zero. To claim a real edge after a thousand trials, you would need a Sharpe comfortably above 1.03, not at it. The full deflated Sharpe ratio goes further, folding in the sample length and the fat tails of real returns as well as the number of trials, and returning something like the probability that the true Sharpe is positive once all of that is accounted for. The practical rule it enforces is blunt: a Sharpe means nothing until you know how hard you searched for it.

The honest number

This is the number that separates a real quant from a backtest tourist. It bakes in the humility the whole course has argued for: your best result is inflated by every strategy you ever tried, whether you admit it or not, and only a Sharpe that clears the haircut, and then survives out-of-sample, deserves any belief. Most impressive backtests do not clear it, which is not a flaw in the method but the truth the method reveals.

What to carry forward

The deflated Sharpe ratio puts a number on the multiple-testing problem by discounting an apparent Sharpe for the number of trials behind it. The expected best from luck grows with the search (about 1.03 for a thousand no-edge tries over a decade of data, exactly the level the search chapter measured), and a discovered Sharpe must clear that haircut, then survive out-of-sample, to be believed. A Sharpe of 1.0 after a thousand tries is effectively zero once deflated. This is the honest number, and it completes the scientific heart of the course. You now know how to find an edge and, far harder, how to tell whether it is real. The next part assumes you have cleared that bar, a rare thing, and builds the survivors into a portfolio.