# Strings, Bytes and Efficient Text Building — Data Structures in Python

Source: https://www.skillbyai.com/en/data-structures-python/s-strings

> Work with immutable strings, bytes and bytearray efficiently.

## Immutable text and mutable buffers

Python **`str`** objects are **immutable sequences of Unicode code points**. Indexing and slicing work like lists, `len` is O(1), and membership (`"pin" in text`) and `find` are O(n·m) in the worst case but fast in practice. Because strings are immutable, every "modification" creates a new string: repeated `s += piece` in a loop can be quadratic, so build text with **`"".join(parts)`**, an `io.StringIO` buffer, or f-strings for small pieces. CPython stores strings compactly using 1, 2 or 4 bytes per character depending on the widest character (PEP 393), so a single emoji can widen a long ASCII string. **`bytes`** are immutable sequences of integers 0–255 for binary data and encoded text (`text.encode("utf-8")`, `data.decode("utf-8")`), and **`bytearray`** is their **mutable** counterpart, useful for building binary buffers. **`memoryview`** gives zero-copy slices of bytes-like objects. String methods (`split`, `strip`, `replace`, `startswith`, `partition`, `casefold`) are implemented in C and are much faster than character-by-character loops, and **`re`** handles patterns.

## Building text and working with bytes

join, StringIO, encoding and a mutable bytearray.

```python
import io

rows = [("o-1", "Pune", 1200), ("o-2", "Delhi", 300)]

# efficient: build a list of pieces and join once
csv_text = "\n".join(f"{oid},{city},{total}" for oid, city, total in rows)
print(csv_text)

buf = io.StringIO()                   # file-like text buffer
for oid, city, total in rows:
    buf.write(f"{oid:<5}{city:>8}{total:>7}\n")
print(buf.getvalue())

name = "नमस्ते"
print(len(name), len(name.encode("utf-8")))     # 6 code points, 18 bytes

packet = bytearray(b"HDR")
packet += (1200).to_bytes(4, "big")              # mutable: append binary data in place
packet[0] = ord("h")
print(bytes(packet))                             # b'hDR\x00\x00\x04\xb0'

view = memoryview(packet)[3:]                    # zero-copy slice
print(int.from_bytes(view, "big"))               # 1200

print("PIN: 411001".partition(": "))             # ('PIN', ': ', '411001')
print("Straße".casefold() == "STRASSE".casefold())   # True: case-insensitive comparison
```

## Use casefold for case-insensitive matching

`lower()` is not enough for all languages (German ß, for example). `casefold()` is designed for caseless comparison of Unicode text.

**Quiz:** Why is "".join(parts) preferred over repeated += for building a long string?

- [x] Strings are immutable, so repeated += may copy the growing string each time, while join computes the size once
- [ ] join sorts the parts
- [ ] join is the only way to concatenate
- [ ] It uses less Unicode

*Answer:* Strings are immutable, so repeated += may copy the growing string each time, while join computes the size once. join allocates the result once and copies each part once.
