Course contents
The unglamorous eighty percent
Raw data is full of splits, bonuses, dividends, gaps, and errors, and using it unadjusted produces nonsense; worse, using tomorrow's information in today's test produces a beautiful lie. This chapter covers adjusting prices for corporate actions and, crucially, point-in-time data that avoids look-ahead bias in fundamentals.
- Adjust prices for splits, bonuses, and dividends correctly
- Explain point-in-time data and how using restated figures creates look-ahead bias
- Treat data cleaning as the foundation, not a chore to rush
The glamorous image of quant work is clever models. The reality is that most of the time goes to cleaning data, and the two most expensive mistakes in all of quant research hide in dirty data: using prices that were not adjusted for corporate actions, and using information you would not have had at the time. This chapter covers both, because getting them wrong is the difference between a backtest that means something and one that only looks like it does.
Adjusting for corporate actions
When a company splits its stock, issues a bonus, or pays a dividend, the quoted price changes without the company's value changing. A one-for-two split turns each share into two, so the quoted price halves overnight. If you feed a backtest the raw quoted prices, it sees a 50% crash that never happened.
# A stock that had a 1-for-2 split: each share became two, so the quoted price
# halved overnight. RAW prices show a fake 50% crash on the split day. A backtest
# using them would believe the stock lost half its value. It did not. The fix is
# to adjust every price before the split, which is what a vendor's "adjusted
# close" already does.
import pandas as pd
raw = pd.DataFrame({
"day": ["Mon", "Tue", "Wed", "Thu", "Fri"],
"close": [1000, 1010, 505, 510, 520], # split takes effect Wed
})
raw["raw_return_pct"] = (raw["close"].pct_change() * 100).round(1)
# Multiply every price BEFORE the split by the split ratio (0.5) to make the
# series continuous.
split_ratio = 0.5
adj = raw["close"].astype(float).copy()
adj.iloc[:2] = adj.iloc[:2] * split_ratio # Mon and Tue were pre-split
raw["adj_close"] = adj
raw["adj_return_pct"] = (raw["adj_close"].pct_change() * 100).round(1)
print(raw.to_string(index=False))
print("\nRaw Wednesday return looks like -50%. Adjusted, it is about 0%.")
print("Always backtest on adjusted prices, never raw quoted prices.")day close raw_return_pct adj_close adj_return_pct Mon 1000 NaN 500.0 NaN Tue 1010 1.0 505.0 1.0 Wed 505 -50.0 505.0 0.0 Thu 510 1.0 510.0 1.0 Fri 520 2.0 520.0 2.0 Raw Wednesday return looks like -50%. Adjusted, it is about 0%. Always backtest on adjusted prices, never raw quoted prices.
The raw series shows Wednesday's return as minus 50%, a catastrophe that exists only in the data, not in reality. Adjusting the prices before the split by the split ratio makes the series continuous, and the true Wednesday return is about zero. Splits, bonuses, and dividends all need this treatment, and a data vendor's "adjusted close" column does it for you. Understand it anyway, because an error in adjustment poisons every single return in your backtest, and a strategy built on poisoned returns is meaningless.
The subtler killer: look-ahead bias
The deadlier trap is look-ahead bias: using information in a backtest that you could not have known at the time. In fundamentals it is everywhere. A company's results for a financial year are not known until they are announced, months after the year ends, and are sometimes restated later still. If your backtest ranks stocks on a year's earnings as though they were known on the first day of that year, it is trading on numbers that did not yet exist, a quiet peek into the future that can make almost any strategy look brilliant.
The fix is point-in-time data: for each date in the backtest, use only the figures that were actually available on that date, lagged for the real reporting delay. Point-in-time data is unglamorous, harder to obtain, and absolutely essential. Using today's restated, complete figures to test a decision you would have made years ago is one of the commonest reasons a backtest that looked wonderful collapses in live trading, and it is entirely self-inflicted.
Why this is the job
It is tempting to rush the data and get to the modelling, and it is a mistake. Cleaning and aligning data honestly is the majority of real quant work, often cited as around eighty percent of it, and it is the foundation every result stands on. Skipping it does not save time. It just moves the failure from your screen, where it is free to fix, to your account, where it is not.
What to carry forward
Most quant work is cleaning data, and two mistakes there are the most expensive in the field. Adjust prices for splits, bonuses, and dividends, or a corporate action masquerades as a 50% crash, as you saw a raw Wednesday return of minus 50% become about zero once adjusted. And use point-in-time data, only what was known on each date, or look-ahead bias quietly lets your test trade on the future and flatters any strategy. This unglamorous work is the foundation of every honest result. The next chapter is the third great data trap, the one hiding in the very choice of which stocks to test on: survivorship bias.