Lesson 7 / 25

Series and DataFrames

Columns with names and types.

A DataFrame is a dict of columns

A Series is a labelled one-dimensional array; a DataFrame is a table of Series sharing an index, each column with its own dtype. Create one from a dict of lists, a list of dicts, files or SQL. Explore with shape, dtypes, head, info and describe, and add columns by vectorised expressions on existing ones. In pandas 3, text columns use a dedicated string dtype (shown as str) instead of generic object.

Labelled tables

pandas adds labels, mixed column types and powerful I/O on top of NumPy.

Three ideas: DataFrames, selection, reading and writing files.
Figure 3.1 — DataFrames, selection and I/O.

Building and exploring a sales table, run

I ran this with Python 3.12.3 and pandas 3.0.6. A five-row DataFrame with text and numeric columns; text columns have the str dtype in pandas 3, and revenue is computed column-wise.

import pandas as pd

df = pd.DataFrame({
    "city": ["Delhi", "Mumbai", "Pune", "Delhi", "Pune"],
    "product": ["pen", "mug", "pen", "lamp", "mug"],
    "units": [10, 4, 7, 1, 3],
    "price": [20.0, 250.0, 20.0, 1499.0, 250.0],
})
print(df)
print(df.shape)
print(df.dtypes)
df["revenue"] = df["units"] * df["price"]
print(df.describe().loc[["mean", "min", "max"], ["units", "revenue"]])

Output:

     city product  units   price
0   Delhi     pen     10    20.0
1  Mumbai     mug      4   250.0
2    Pune     pen      7    20.0
3   Delhi    lamp      1  1499.0
4    Pune     mug      3   250.0
(5, 4)
city           str
product        str
units        int64
price      float64
dtype: object
      units  revenue
mean    5.0    717.8
min     1.0    140.0
max    10.0   1499.0

Look before you analyse

Check shape, dtypes and a few rows first; wrong dtypes (numbers stored as text) cause most early bugs.

Quick check: What does each DataFrame column have?

  • Its own dtype
  • The same dtype as all other columns
  • No dtype
  • A separate index per row
Answer

Its own dtype — Columns are typed Series.