पाठ 9 / 25

Reading and Writing Data

CSV and beyond.

read_csv options that matter

pd.read_csv handles most text data: parse_dates converts date columns, usecols loads only needed columns, dtype fixes types, and empty fields become NaN. Write with to_csv(index=False) to avoid an extra index column. For larger or typed data, Parquet (read_parquet/to_parquet) is faster and preserves dtypes; read_sql reads from databases.

Reading a CSV with dates and missing values, run

I ran this with Python 3.12.3 and pandas 3.0.6. The CSV is read from an in-memory string: the date column is parsed as datetime, the empty amount becomes NaN, and the round trip writes 120.5 without the index.

import io
import pandas as pd

csv_text = """order_id,date,amount,status
1,2026-01-05,120.50,paid
2,2026-01-06,,refunded
3,2026-01-06,75.00,paid
"""
df = pd.read_csv(io.StringIO(csv_text), parse_dates=["date"])
print(df)
print(df.dtypes)
out = io.StringIO()
df.to_csv(out, index=False)
print(out.getvalue().splitlines()[1])
print(pd.read_csv(io.StringIO(csv_text), usecols=["order_id", "amount"]).shape)

Output:

   order_id       date  amount    status
0         1 2026-01-05   120.5      paid
1         2 2026-01-06     NaN  refunded
2         3 2026-01-06    75.0      paid
order_id             int64
date        datetime64[us]
amount             float64
status                 str
dtype: object
1,2026-01-05,120.5,paid
(3, 2)

Use Parquet for intermediate data

Parquet keeps dtypes and is much smaller and faster than CSV for data passed between steps.

त्वरित जाँच: What does an empty numeric field in a CSV become after read_csv?

  • An empty string in a float column
  • 0
  • NaN
  • An error
Answer

NaN — Missing values become NaN.