Skip to content
Course contents
Working with data

Trusting your data first

Before computing anything, you inspect and clean the data, checking its size, its types, and its missing values, because a wrong result from dirty data is worse than no result. Real prices have gaps, splits, and errors.

7 min readChapter 18 of 30
What you will learn
  • Inspect a DataFrame with head, tail, info, and describe
  • Find and handle missing values
  • Reason about the real-world messiness of price data

There is a rule in data work that beginners learn the hard way: a confident, precise result computed from dirty data is worse than no result at all, because you will trust it. Before you compute a single return or draw a single chart, you inspect the data and clean it, so that what follows rests on something solid. Real market data is messy, gaps on holidays, a halted stock, a split that makes a price look absurd, a value that arrived as text, and a few minutes of checking saves you from hours of chasing a wrong number.

Look before you compute

Before computing anything, inspect and clean the data: check its shape and types and handle missing values, because a wrong result from dirty data is worse than none.
Before computing anything, inspect and clean the data: check its shape and types and handle missing values, because a wrong result from dirty data is worse than none.

pandas gives you quick ways to inspect a DataFrame, and you use them every time you load data. This program checks the sample before doing anything with it.

ExampleInspecting shape, types, and missing values, then handling a gapch18/cleaning.py
# Inspect and clean data before trusting any calculation on it.
import pandas as pd

df = pd.read_csv("sample_prices.csv", parse_dates=["Date"], index_col="Date")

print("Rows, columns:", df.shape)
print("Column types:")
print(df.dtypes)
print("Missing values per column:")
print(df.isna().sum())

# Demonstrate handling a gap: make one Close missing, then fill it.
df.loc[df.index[5], "Close"] = None
print("After introducing one gap, missing Closes:", int(df["Close"].isna().sum()))
df["Close"] = df["Close"].ffill()        # carry the previous value forward
print("After forward-fill, missing Closes:", int(df["Close"].isna().sum()))
Output
Rows, columns: (180, 5)
Column types:
Open      float64
High      float64
Low       float64
Close     float64
Volume      int64
dtype: object
Missing values per column:
Open      0
High      0
Low       0
Close     0
Volume    0
dtype: int64
After introducing one gap, missing Closes: 1
After forward-fill, missing Closes: 0

The output walks through the essential checks. df.shape confirms 180 rows and 5 columns, the size you expected. df.dtypes shows the type of each column, the prices as float64 and volume as int64, which matters because a price column accidentally loaded as text would refuse to do arithmetic. df.isna().sum() counts the missing values in each column, here all zero, so this data is complete. Alongside these, two more inspectors are worth knowing: df.info() summarises the columns, types, and counts in one call, and df.describe() gives quick statistics, the mean, minimum, maximum, and quartiles of each column, which often exposes an absurd value at a glance.

Handling what is missing

Real data is rarely so clean, so the second half of the program shows what to do about a gap. It deliberately sets one Close to missing, and isna().sum() then reports one missing Close. The fix here is forward-fill, ffill, which carries the previous day's value into the gap, a common and sensible choice for a missing price, after which the count of missing values is zero again. The other common response is dropna, which removes rows with missing values entirely, appropriate when a gap cannot sensibly be filled. Which you choose depends on the situation, but the discipline is the same: find the missing values, decide deliberately how to handle them, and never let them slip silently into a calculation, where they spread and corrupt the result.

The deeper habit is skepticism toward your own data. Before trusting a number your analysis produces, you should already have checked that the data it came from is the right size, the right types, and free of gaps you did not account for. Analysts spend more time cleaning and checking data than computing on it, and that is not wasted time; it is what makes the computing trustworthy.

What to carry forward

Before computing anything, inspect the data, its shape, its column types, its summary statistics, and its missing values, and handle any gaps deliberately with forward-fill or by dropping rows, never letting them slip silently into a calculation. Real prices are messy, and a confident result from dirty data is the most dangerous kind, so cleaning and checking is time that makes everything after it trustworthy.

With clean, trusted data in hand, you can start asking it questions. The next chapter selects and filters the table, pulling out the columns and the rows that answer a specific market question.