# Reading and Writing Data — NumPy / Pandas / scikit-learn

Source: https://www.skillbyai.com/en/numpy-pandas-sklearn/p-io

> CSV and beyond.

## read_csv options that matter

`pd.read_csv` handles most text data: `parse_dates` converts date columns, `usecols` loads only needed columns, `dtype` fixes types, and empty fields become NaN. Write with `to_csv(index=False)` to avoid an extra index column. For larger or typed data, **Parquet** (`read_parquet`/`to_parquet`) is faster and preserves dtypes; `read_sql` reads from databases.

## Reading a CSV with dates and missing values, run

I ran this with Python 3.12.3 and pandas 3.0.6. The CSV is read from an in-memory string: the date column is parsed as datetime, the empty amount becomes NaN, and the round trip writes 120.5 without the index.

```python
import io
import pandas as pd

csv_text = """order_id,date,amount,status
1,2026-01-05,120.50,paid
2,2026-01-06,,refunded
3,2026-01-06,75.00,paid
"""
df = pd.read_csv(io.StringIO(csv_text), parse_dates=["date"])
print(df)
print(df.dtypes)
out = io.StringIO()
df.to_csv(out, index=False)
print(out.getvalue().splitlines()[1])
print(pd.read_csv(io.StringIO(csv_text), usecols=["order_id", "amount"]).shape)
```

Output:

```
   order_id       date  amount    status
0         1 2026-01-05   120.5      paid
1         2 2026-01-06     NaN  refunded
2         3 2026-01-06    75.0      paid
order_id             int64
date        datetime64[us]
amount             float64
status                 str
dtype: object
1,2026-01-05,120.5,paid
(3, 2)
```

## Use Parquet for intermediate data

Parquet keeps dtypes and is much smaller and faster than CSV for data passed between steps.

**Quiz:** What does an empty numeric field in a CSV become after read_csv?

- [ ] An empty string in a float column
- [ ] 0
- [x] NaN
- [ ] An error

*Answer:* NaN. Missing values become NaN.
