पाठ 2 / 26
Images as Arrays of Pixels
Height, width, channels.
Numbers in a grid
A digital image is a grid of pixels. A grayscale image is a 2D array (height x width) with one brightness value per pixel; a colour image adds a channel dimension, usually three channels for red, green and blue. Values are typically 8-bit integers from 0 to 255, or floats scaled to 0 to 1 for neural networks. Libraries use different layouts: NumPy, OpenCV and scikit-image use (height, width, channels); PyTorch uses (channels, height, width). Many bugs come from mixing layouts or value ranges.
Inspecting a photo as an array, run
I ran this on CPU with Python 3, OpenCV 5.0.0, scikit-image 0.26.0, PyTorch 2.14.1 and torchvision 0.29.1, using scikit-image's bundled sample photos and torchvision's published pretrained weights. The bundled astronaut photo is a 512 x 512 x 3 array of 8-bit values from 0 to 255; the mean red value (141.6) is higher than green and blue. The camera photo is grayscale: one value per pixel.
from skimage import data
img = data.astronaut() # a bundled RGB photo
print("shape (height, width, channels):", img.shape, "| dtype:", img.dtype)
print("top-left pixel RGB:", img[0, 0].tolist())
print("value range:", img.min(), "to", img.max())
print("mean per channel (R, G, B):", img.reshape(-1, 3).mean(axis=0).round(1).tolist())
gray = data.camera()
print("grayscale image shape:", gray.shape, "-> one number per pixel")
Output:
shape (height, width, channels): (512, 512, 3) | dtype: uint8 top-left pixel RGB: [154, 147, 151] value range: 0 to 255 mean per channel (R, G, B): [141.6, 105.8, 96.5] grayscale image shape: (512, 512) -> one number per pixel
Print shape, dtype and range
Before any processing, print the array shape, data type and min/max; it catches most layout and scaling bugs.
त्वरित जाँच: What layout does PyTorch expect for an image tensor?
- (channels, height, width)
- (height, width, channels)
- (width, height)
- (pixels, labels)
Answer
(channels, height, width) — Convert with permute when moving from NumPy.