Course contents
Testing on data you have never seen
The defence against a false edge is to judge a strategy only on data it was never tuned on, using held-out test sets and walk-forward analysis that repeatedly trains on the past and tests on the next unseen slice. This chapter builds that validation, extending the walk-forward idea from Algorithmic Trading, and names the leakage mistakes that quietly reintroduce cheating.
- Split data into training, validation, and test sets and apply walk-forward analysis
- Identify and avoid data leakage that reintroduces look-ahead
- Treat out-of-sample performance as the only performance that counts
If searching many strategies manufactures a false winner, how do you catch it? With the oldest defence in science: test your chosen strategy on data it has never seen. A strategy that was picked for its luck on one set of data has no reason to be lucky on a fresh set, so out-of-sample testing is where false edges go to die. This chapter builds that defence and its stronger cousin, walk-forward analysis.
The held-out test
The core move is to hold data back. Split your history into an in-sample period, where you research, tune, and search as much as you like, and an out-of-sample period, which you do not touch until the very end and use only once, to judge the single strategy you finally chose. If the edge is real it shows up in both periods; if it was search-luck, it evaporates the moment it meets data it was not fitted to. Watch it happen.
# Out-of-sample testing is the defense against the multiple-testing trap. We search
# many no-edge strategies, pick the one with the best Sharpe on the IN-SAMPLE half,
# then see how it does on the OUT-OF-SAMPLE half it was never chosen on. The luck
# does not repeat.
import numpy as np
rng = np.random.default_rng(8)
n = 240 # 20 years of monthly returns
half = n // 2
def sharpe(x):
return x.mean() / x.std(ddof=1) * np.sqrt(12) if x.std(ddof=1) > 0 else 0.0
trials = 2000
strategies = rng.normal(0, 0.04, (trials, n)) # all no-edge
in_sharpes = np.array([sharpe(strategies[i, :half]) for i in range(trials)])
best = int(in_sharpes.argmax())
oos = sharpe(strategies[best, half:])
avg_oos = np.mean([sharpe(strategies[i, half:]) for i in range(trials)])
print(f"Searched {trials} no-edge strategies over {n} months.")
print(f"Best IN-SAMPLE Sharpe: {in_sharpes[best]:5.2f} (looks excellent, pure luck)")
print(f"Its OUT-OF-SAMPLE Sharpe: {oos:5.2f} (the luck did not repeat)")
print(f"Average OOS Sharpe of all: {avg_oos:5.2f} (as expected for no edge: about zero)")
print("\nA strategy chosen for its in-sample luck has no reason to do well on new")
print("data. Only out-of-sample performance is honest evidence. Walk-forward testing")
print("repeats this, always training on the past and testing on the next unseen slice.")Searched 2000 no-edge strategies over 240 months. Best IN-SAMPLE Sharpe: 1.16 (looks excellent, pure luck) Its OUT-OF-SAMPLE Sharpe: 0.42 (the luck did not repeat) Average OOS Sharpe of all: -0.00 (as expected for no edge: about zero) A strategy chosen for its in-sample luck has no reason to do well on new data. Only out-of-sample performance is honest evidence. Walk-forward testing repeats this, always training on the past and testing on the next unseen slice.
We search two thousand no-edge strategies and pick the one with the best Sharpe on the in-sample half, which comes out at a convincing 1.16. Then we test that exact strategy on the out-of-sample half it was never chosen on, and its Sharpe falls to 0.42, while the average out-of-sample Sharpe across all the strategies is essentially zero, exactly as it should be for strategies with no edge. The in-sample star was pure luck, and the out-of-sample test caught it. This is the direct antidote to the previous chapter's trap.
Walk-forward, and the leakage that undoes it
A single split spends the out-of-sample data only once, which is a little wasteful and a little fragile. Walk-forward analysis does better: it repeatedly trains on a window of past data and tests on the next unseen window, then rolls forward and does it again, marching through history. You get many out-of-sample tests instead of one, and you see whether the edge holds across different periods rather than just surviving a single lucky slice. This is the same walk-forward you met in Algorithmic Trading, now the backbone of honest quant validation.
All of it rests on one discipline: never let the future leak in. Data leakage is any way information from the test period sneaks into the training, and it is subtle. Scaling your data using statistics computed over the whole history, including the test period, leaks. Choosing a universe or a threshold by glancing at how it does on the test set leaks. Retesting on the out-of-sample data after a poor result, then tuning, quietly turns it into in-sample data. The rule is absolute: the moment you use the out-of-sample data to make any choice, it stops being out-of-sample, and your defence is gone.
What to carry forward
Out-of-sample testing is the defence against the multiple-testing trap: judge a chosen strategy only on data it was never tuned on, and a lucky in-sample winner collapses, as you saw a 1.16 Sharpe fall to 0.42. Walk-forward analysis strengthens this by rolling through history, training on the past and testing on the next unseen slice repeatedly. And the whole thing depends on avoiding data leakage, because the moment the test data influences any choice it stops being a fair test. Out-of-sample testing tells you whether an edge survives; the next chapter puts a precise number on how much to discount it for the search, with the deflated Sharpe ratio.