Skip to content
Course contents
Putting it together

Writing code you can trust

Code that touches money must be trustworthy, so this chapter covers the habits, readable code, reproducibility, checking your data and your assumptions, and the deeper pitfalls, survivorship bias and curve-fitting, that fool even careful people. The goal is honest, reliable analysis.

9 min readChapter 29 of 30
What you will learn
  • Apply basic code-hygiene and reproducibility habits
  • Explain survivorship bias and curve-fitting and how to guard against them
  • Frame a backtest as a hypothesis to doubt, not a promise to believe

The honest backtest a couple of chapters ago showed the crossover losing money. A determined beginner, unwilling to accept that, can make almost any losing strategy look like a winner, not by cheating on purpose, but by falling into traps that fool careful people every day. This chapter is about those traps, and about the habits that keep your analysis honest. It is the most important chapter in the course for anyone who intends to test their own ideas.

Curve-fitting, demonstrated

Curve-fitting fits the noise in the past perfectly and fails in the future; a simpler model that fits less can generalise better. Illustrative.
Curve-fitting fits the noise in the past perfectly and fails in the future; a simpler model that fits less can generalise better. Illustrative.

The most seductive trap is curve-fitting, also called overfitting: tuning a strategy's settings until it fits the past beautifully, capturing the noise of one particular history rather than any real pattern. Here is how easily it happens.

ExampleTrying many parameter combinations and keeping the best, which is curve-fittingch29/overfitting.py
# Trying many parameter combinations on one dataset is curve-fitting.
import itertools
import pandas as pd

df = pd.read_csv("sample_prices.csv", parse_dates=["Date"], index_col="Date")
ret = df["Close"].pct_change().fillna(0)

best = None
for fast, slow in itertools.product([5, 10, 20], [30, 50, 100]):
    if fast >= slow:
        continue
    signal = (df["Close"].rolling(fast).mean() > df["Close"].rolling(slow).mean()).astype(int)
    position = signal.shift(1).fillna(0)
    total = (1 + position * ret).prod() - 1
    if best is None or total > best[2]:
        best = (fast, slow, total)

print(f"Best combo on THIS sample: fast={best[0]}, slow={best[1]} -> {best[2] * 100:.2f}%")
print("It was chosen by trying many on one dataset. That is curve-fitting,")
print("and its future performance is almost always worse than this number.")
Output
Best combo on THIS sample: fast=10, slow=30 -> 6.37%
It was chosen by trying many on one dataset. That is curve-fitting,
and its future performance is almost always worse than this number.

The honest crossover, with its 20 and 50-day averages, lost 8.11%. But this program tries several combinations of fast and slow windows and keeps whichever did best on the sample, and it finds that a 10-and-30-day version would have returned plus 6.37%. It is tempting to announce that as the strategy. Do not. That plus 6.37% was chosen precisely because it was the luckiest fit to this one dataset; the search itself manufactured the number. Run these combinations on a different period and the "best" one will very likely disappoint, because you fitted the parameters to the past's noise, not to a durable edge. The more combinations you try, the more certainly you are curve-fitting, and the better the backtest looks, the more suspicious you should be.

The guard against it is discipline: prefer few parameters over many, be suspicious of a result that required searching to find, and, above all, test a chosen strategy on data you did not use to build it, so-called out-of-sample testing. A strategy that only shines on the data it was tuned to is not a strategy; it is a memory.

The other pitfalls

A few more traps deserve naming, because they defeat even careful analysts.

Survivorship bias creeps in when your data includes only the companies or funds that still exist today. Test a strategy on the current members of an index and you have quietly excluded every company that went bankrupt or was removed, which flatters any result, because you tested only on survivors. Real history includes the failures, and honest data must too.

Lookahead bias, from the backtest chapter, is worth repeating: any use of information you could not have had at the time, a signal computed from a price you had not yet seen, a figure revised after the fact, inflates results and must be hunted down.

Ignoring costs and slippage flatters a strategy, especially an active one, and small samples mislead: 180 days, like our sample, is far too little to conclude anything, and a real study needs years across different market conditions.

Practices that keep you honest

Against all this stand a few plain habits. Write readable code with clear names, so you and others can check it, because a bug in a backtest usually helps you, by inflating the result, and so goes unnoticed. Make your work reproducible: fix any random seed, note the versions of your tools, and keep the data, so a result can be reproduced rather than merely believed. Check your data and your assumptions before trusting an output. And hold the attitude this whole part has built toward: a backtest is a hypothesis to be doubted and stress-tested, never a promise to be believed. The more a result excites you, the harder you should try to break it.

What to carry forward

Code that touches money must be trustworthy, and the traps are real: curve-fitting manufactures a winner from noise (a search turned our losing rule into a +6.37% fantasy), survivorship bias tests only on survivors, and lookahead, costs, and tiny samples all flatter results. The habits that protect you are few parameters, out-of-sample testing, reproducible code, and the settled attitude that a backtest is a hypothesis to doubt, hardest when it excites you most.

That is the honest heart of quantitative work. The final chapter recaps the whole workflow and looks ahead to algorithmic trading, where analysis becomes live execution, a larger and harder step than it first appears.