पाठ 10 / 25
Handling Missing Values
Count, drop or fill.
isna, dropna, fillna
Start with df.isna().sum() to see missing values per column. dropna() removes rows with any missing value (often too aggressive); dropna(subset=[...]) targets required columns. fillna replaces values: a median for numbers, a label such as "unknown" for categories. Decide per column, record the choice, and in machine learning fit imputers on training data only.
From messy to reliable
Missing values, wrong types and accidental copies are the main sources of bad analyses.
Inspecting and filling missing values, run
I ran this with Python 3.12.3 and pandas 3.0.6. Plain dropna keeps only one complete row; filling age with the median (28.0) and city with "unknown" keeps all four, while dropping only rows without a name keeps three.
import numpy as np
import pandas as pd
df = pd.DataFrame({
"name": ["Asha", "Ben", None, "Dev"],
"age": [31, np.nan, 25, np.nan],
"city": ["Pune", "Delhi", "Pune", None],
})
print(df.isna().sum())
print(df.dropna())
filled = df.assign(
age=df["age"].fillna(df["age"].median()),
city=df["city"].fillna("unknown"),
)
print(filled)
print(df.dropna(subset=["name"]).shape)
Output:
name 1 age 2 city 1 dtype: int64 name age city 0 Asha 31.0 Pune name age city 0 Asha 31.0 Pune 1 Ben 28.0 Delhi 2 NaN 25.0 Pune 3 Dev 28.0 unknown (3, 3)
Ask why data is missing
Missing values are often informative (a skipped form field, a failed sensor); a "was_missing" flag column can help models.
त्वरित जाँच: What does df.dropna() do by default?
- Drops every row that has any missing value
- Fills missing values with 0
- Drops columns only
- Nothing unless inplace=True
Answer
Drops every row that has any missing value — Often too aggressive; use subset.