# Reserving and Controlling Output — Prompt & Context Engineering

Source: https://www.skillbyai.com/en/prompt-context-engineering/b-output

> The answer needs room and a shape.

## Max tokens, length guidance and stop points

Set a **maximum output tokens** value that fits the longest valid answer, and keep that space free in the window. If answers are cut off mid-sentence, either the limit is too low or the prompt invites rambling. Give **explicit length guidance** (for example, three bullet points, or under 120 words) and a **format** so the model knows when it is done. Check the API's stop reason: a response that ended because it hit the length limit is incomplete and should not be treated as a final answer, especially for structured output.

## Handling the stop reason (sketch)

Field names differ by provider; the logic is what matters. Not run here.

```python
resp = call_model(prompt, max_output_tokens=600)
if resp.stop_reason == "length":        # hit the output limit
    log.warning("truncated answer", request_id=resp.id)
    resp = call_model(prompt + "\nAnswer in at most 5 bullet points.", max_output_tokens=600)
answer = resp.text
```

## Track truncation rate

Count how often answers stop on the length limit; a rising rate is an early warning that prompts or inputs have grown.

**Quiz:** A response stopped because it hit the length limit. What is true?

- [ ] It means the input was empty
- [ ] It is always complete
- [ ] It means the model refused
- [x] It may be incomplete and should be handled, not used as final

*Answer:* It may be incomplete and should be handled, not used as final. Check the stop reason before trusting the output.
