Lesson 6 / 25

Reserving and Controlling Output

The answer needs room and a shape.

Max tokens, length guidance and stop points

Set a maximum output tokens value that fits the longest valid answer, and keep that space free in the window. If answers are cut off mid-sentence, either the limit is too low or the prompt invites rambling. Give explicit length guidance (for example, three bullet points, or under 120 words) and a format so the model knows when it is done. Check the API's stop reason: a response that ended because it hit the length limit is incomplete and should not be treated as a final answer, especially for structured output.

Handling the stop reason (sketch)

Field names differ by provider; the logic is what matters. Not run here.

resp = call_model(prompt, max_output_tokens=600)
if resp.stop_reason == "length":        # hit the output limit
    log.warning("truncated answer", request_id=resp.id)
    resp = call_model(prompt + "\nAnswer in at most 5 bullet points.", max_output_tokens=600)
answer = resp.text

Track truncation rate

Count how often answers stop on the length limit; a rising rate is an early warning that prompts or inputs have grown.

Quick check: A response stopped because it hit the length limit. What is true?

  • It means the input was empty
  • It is always complete
  • It means the model refused
  • It may be incomplete and should be handled, not used as final
Answer

It may be incomplete and should be handled, not used as final — Check the stop reason before trusting the output.