Skip to content
Course contents
Data, the raw material

The raw material

Quant strategies run on data: prices and volumes, fundamentals, corporate actions, and a defined universe of instruments, each with its own quality and pitfalls. This chapter surveys what data a quant needs, where Indian market data comes from, and why the quality of the data sets a hard ceiling on the quality of any result.

9 min readChapter 5 of 26
What you will learn
  • List the data types a quant strategy needs (price, volume, fundamentals, corporate actions, universe)
  • Identify sources of Indian market and fundamental data
  • Explain why data quality caps result quality

Every quant strategy is only as good as the data under it, and data is where most of the real work, and most of the hidden failure, lives. Before any strategy, a quant assembles the raw material: what the market did, what companies are worth, and which instruments are even in play. This chapter surveys that raw material, the shape it takes, and the plain truth that no amount of clever modelling can rescue a result built on bad data.

The kinds of data

A quant needs several kinds of data: prices, company fundamentals, corporate actions, and the changing membership of a universe.
A quant needs several kinds of data: prices, company fundamentals, corporate actions, and the changing membership of a universe.

A quant draws on a few distinct kinds of data, each with its own uses and traps.

Market data is the backbone: the open, high, low, close, and volume for each instrument, the OHLCV you met in Python for Trading. For quant work you need it not for one stock but for the whole universe, over long histories, and adjusted for corporate actions.

Fundamental data is the company numbers from Fundamental Analysis: earnings, book value, debt, cash flow, and the ratios built from them like the price-to-earnings ratio and return on equity. These drive the value and quality factors of Part 3.

Corporate actions are the events that change a raw price without changing value: splits, bonuses, and dividends. Handle them wrongly and your returns are corrupted, as the next chapter shows.

The universe is the set of instruments you consider trading, and defining it, which stocks, filtered how, is itself a design choice with real consequences, as the chapter after next shows.

The shape of quant data

Where Python for Trading worked with one stock's price series, a quant works with a panel: a table with one row per stock and columns for the characteristics a strategy ranks on. Here is that shape.

ExampleThe panel a quant ranks: one row per stock, columns to rank onch05/quant_data_shapes.py
# The raw material of quant work is a panel: one row per stock, with the columns a
# strategy ranks on (price, size, valuation, quality, past return). This is an
# illustrative snapshot of a few NSE-listed stocks. Real data covers the whole
# universe over many years; this is the shape of it. All figures illustrative.
import pandas as pd

stocks = pd.DataFrame({
    "symbol":      ["RELIANCE", "TCS", "HDFCBANK", "INFY", "ITC", "TATASTEEL"],
    "sector":      ["Energy", "IT", "Banking", "IT", "FMCG", "Metals"],
    "price":       [1400, 3900, 1650, 1500, 460, 150],
    "mktcap_cr":   [1900000, 1400000, 1250000, 620000, 570000, 190000],
    "pe":          [24.0, 30.0, 18.0, 26.0, 22.0, 12.0],
    "roe_pct":     [9.0, 47.0, 17.0, 31.0, 28.0, 8.0],
    "ret_12m_pct": [12.0, 8.0, 5.0, -3.0, 15.0, 25.0],
})

print("The panel a quant works with (one row per stock):")
print(stocks.to_string(index=False))

print("\nA factor is just a ranking. Cheapest by P/E (a value tilt):")
print(stocks.sort_values("pe")[["symbol", "pe"]].head(3).to_string(index=False))

print("\nHighest quality by return on equity:")
print(stocks.sort_values("roe_pct", ascending=False)[["symbol", "roe_pct"]].head(3).to_string(index=False))
Output
The panel a quant works with (one row per stock):
   symbol  sector  price  mktcap_cr   pe  roe_pct  ret_12m_pct
 RELIANCE  Energy   1400    1900000 24.0      9.0         12.0
      TCS      IT   3900    1400000 30.0     47.0          8.0
 HDFCBANK Banking   1650    1250000 18.0     17.0          5.0
     INFY      IT   1500     620000 26.0     31.0         -3.0
      ITC    FMCG    460     570000 22.0     28.0         15.0
TATASTEEL  Metals    150     190000 12.0      8.0         25.0

A factor is just a ranking. Cheapest by P/E (a value tilt):
   symbol   pe
TATASTEEL 12.0
 HDFCBANK 18.0
      ITC 22.0

Highest quality by return on equity:
symbol  roe_pct
   TCS     47.0
  INFY     31.0
   ITC     28.0

This is the raw material of everything in Part 3. A factor, the central quant idea, is simply a ranking of this panel on one of its columns: rank by the price-to-earnings ratio and you have a value tilt, rank by return on equity and you have a quality tilt. Real data extends this panel to the entire universe across many years, a large table indexed by both date and stock, but the shape is the same, and the strategy is a rule for ranking it.

India's data, and why quality is everything

For Indian markets, prices come from the exchanges through a broker's own SDK or an open library like yfinance, as settled in the earlier coding courses; fundamentals come from company filings and data vendors; corporate actions come from the exchanges. The sources matter less than one hard rule that governs all of them: data quality sets a ceiling on result quality. Garbage in, garbage out is not a slogan here, it is the single most common cause of a backtest that looks wonderful and fails live. A dataset contaminated by the biases of the next two chapters, survivorship and look-ahead, will produce a beautiful and completely false result, no matter how clever the model on top of it. That is why a serious quant spends more time on data than on strategies.

What to carry forward

A quant assembles market data, fundamentals, corporate actions, and a defined universe into a panel, one row per stock, and a factor is just a ranking of that panel on a column, which is the raw material of every strategy in Part 3. Indian prices come from a broker SDK or an open library, fundamentals from filings and vendors, but the sources matter less than the rule that data quality caps result quality: a contaminated dataset produces a false result whatever the model. The next two chapters are the two biggest contaminations, and the first is the everyday one every backtest must handle: cleaning and adjusting the data.