# Types, Strings and Categories — NumPy / Pandas / scikit-learn

Source: https://www.skillbyai.com/en/numpy-pandas-sklearn/k-types

> Make columns mean what they say.

## to_numeric, .str and Categorical

Convert text numbers with `pd.to_numeric(errors="coerce")` so bad values become NaN instead of crashing. The `.str` accessor applies vectorised string operations (strip, upper, contains, startswith). **Categorical** columns save memory for repeated labels and can be **ordered** so sorting follows S < M < L rather than the alphabet. Convert dates with `pd.to_datetime`.

## Cleaning SKUs, prices and sizes, run

I ran this with Python 3.12.3 and pandas 3.0.6. SKUs are trimmed and upper-cased, "n/a" becomes NaN, and the ordered size category sorts as S, M, L.

```python
import pandas as pd

df = pd.DataFrame({
    "sku": [" pen-01", "MUG-02 ", "pen-03"],
    "price": ["20", "250", "n/a"],
    "size": ["S", "L", "M"],
})
df["sku"] = df["sku"].str.strip().str.upper()
df["price"] = pd.to_numeric(df["price"], errors="coerce")
df["size"] = pd.Categorical(df["size"], categories=["S", "M", "L"], ordered=True)
print(df)
print(df.dtypes)
print(df["sku"].str.startswith("PEN").tolist())
print(df.sort_values("size")["size"].tolist())
```

Output:

```
      sku  price size
0  PEN-01   20.0    S
1  MUG-02  250.0    L
2  PEN-03    NaN    M
sku           str
price     float64
size     category
dtype: object
[True, False, True]
['S', 'M', 'L']
```

## Validate after cleaning

Assert expectations (no negative prices, known categories) after cleaning so bad inputs fail loudly.

**Quiz:** What does errors="coerce" do in pd.to_numeric?

- [ ] Converts to strings
- [ ] Raises an error for bad values
- [ ] Deletes the column
- [x] Turns unparseable values into NaN

*Answer:* Turns unparseable values into NaN. Bad values become missing.
