Lesson 6 / 25
Strings, Bytes and Efficient Text Building
Work with immutable strings, bytes and bytearray efficiently.
Immutable text and mutable buffers
Python str objects are immutable sequences of Unicode code points. Indexing and slicing work like lists, len is O(1), and membership ("pin" in text) and find are O(n·m) in the worst case but fast in practice. Because strings are immutable, every "modification" creates a new string: repeated s += piece in a loop can be quadratic, so build text with "".join(parts), an io.StringIO buffer, or f-strings for small pieces. CPython stores strings compactly using 1, 2 or 4 bytes per character depending on the widest character (PEP 393), so a single emoji can widen a long ASCII string. bytes are immutable sequences of integers 0–255 for binary data and encoded text (text.encode("utf-8"), data.decode("utf-8")), and bytearray is their mutable counterpart, useful for building binary buffers. memoryview gives zero-copy slices of bytes-like objects. String methods (split, strip, replace, startswith, partition, casefold) are implemented in C and are much faster than character-by-character loops, and re handles patterns.
Building text and working with bytes
join, StringIO, encoding and a mutable bytearray.
import io
rows = [("o-1", "Pune", 1200), ("o-2", "Delhi", 300)]
# efficient: build a list of pieces and join once
csv_text = "\n".join(f"{oid},{city},{total}" for oid, city, total in rows)
print(csv_text)
buf = io.StringIO() # file-like text buffer
for oid, city, total in rows:
buf.write(f"{oid:<5}{city:>8}{total:>7}\n")
print(buf.getvalue())
name = "नमस्ते"
print(len(name), len(name.encode("utf-8"))) # 6 code points, 18 bytes
packet = bytearray(b"HDR")
packet += (1200).to_bytes(4, "big") # mutable: append binary data in place
packet[0] = ord("h")
print(bytes(packet)) # b'hDR\x00\x00\x04\xb0'
view = memoryview(packet)[3:] # zero-copy slice
print(int.from_bytes(view, "big")) # 1200
print("PIN: 411001".partition(": ")) # ('PIN', ': ', '411001')
print("Straße".casefold() == "STRASSE".casefold()) # True: case-insensitive comparisonUse casefold for case-insensitive matching
lower() is not enough for all languages (German ß, for example). casefold() is designed for caseless comparison of Unicode text.
Quick check: Why is "".join(parts) preferred over repeated += for building a long string?
- Strings are immutable, so repeated += may copy the growing string each time, while join computes the size once
- join sorts the parts
- join is the only way to concatenate
- It uses less Unicode
Answer
Strings are immutable, so repeated += may copy the growing string each time, while join computes the size once — join allocates the result once and copies each part once.