Lesson 2 / 25
Tokens and the Context Window
Everything is counted in tokens, including the answer.
The window is a hard limit
Models read and write tokens, pieces of words. The context window is the maximum number of tokens for input plus output on one call; anything beyond it is rejected or cut. In English a token is often about four characters, but code, numbers and other languages (including Hindi) can use many more tokens per word. Price and latency also scale with tokens. For budgeting, a rough estimate is fine; for limits, use the model's own tokenizer or the token count returned by the API. Large windows do not mean you should fill them: long inputs cost more and can make the model miss details.
Two rough token estimates, run
I ran this with plain Python 3 (standard library only); the data is made-up example data. An 83-character sentence of 15 words gives estimates of 21 and 20 tokens. These rules of thumb are for planning only; the real tokenizer decides.
text = "Context engineering decides what the model sees, in what order, and how much of it."
words = len(text.split())
chars = len(text)
print("words:", words, "| characters:", chars)
print("rough token estimate (chars / 4):", round(chars / 4))
print("rough token estimate (words * 1.3):", round(words * 1.3))
print("use the real tokenizer of your model for exact counts")
Output:
words: 15 | characters: 83 rough token estimate (chars / 4): 21 rough token estimate (words * 1.3): 20 use the real tokenizer of your model for exact counts
Measure Hindi separately
If you serve several languages, count tokens on real samples of each; the same message can cost several times more in some scripts.
Quick check: What counts against the context window?
- Input tokens and the generated output tokens
- Only the user's question
- Only the system prompt
- Only images
Answer
Input tokens and the generated output tokens — Reserve room for the answer as well.