The budget equation
For a model whose context allowance includes the input and generated tokens, reserve the answer before packing in retrieval:
fixed input = instructions + history + current message + tool data
available for retrieval = context window - fixed input - output reserve - margin
maximum chunks = max(0, floor(available for retrieval / tokens per chunk))
The output reserve includes reasoning when your endpoint counts it against that allowance. Confirm the exact model's separate input/output caps and reasoning rules. A request can satisfy this arithmetic and still violate a provider-specific limit.
Worked example: why ten chunks fit but eleven do not
Assume a 16,000-token window. These are teaching numbers, not a claim about a particular provider or model:
| Allocation | Tokens |
|---|---|
| System instructions | 800 |
| Conversation history | 2,500 |
| Current message | 200 |
| Tool definitions and results | 500 |
| Output + reasoning reserve | 3,000 |
| Safety margin | 1,000 |
| Available for retrieved content | 8,000 |
At 800 tokens per chunk, 10 chunks fit exactly alongside the selected reserves. Eleven chunks require 8,800 tokens, exceeding the total plan by 800. Eight chunks leave 1,600 tokens of additional headroom. Try all three values in the calculator.
If chunk sizes vary, use a conservative per-chunk bound for this planning model. For precise packing, count each selected chunk with its labels and separators, assemble the complete request and count again. An average of 800 does not ensure that every group of ten will fit.
Five mistakes this arithmetic makes visible
- Spending the entire window on input. A large prompt still needs room for the response. Output limits and context limits are different constraints.
- Forgetting tools. Schemas, tool outputs and message wrappers can take substantial space. Tool use is not context-free.
- Counting words instead of model tokens. Tokenization varies by model, language and content. Do not treat a fixed characters-per-token rule as exact.
- Assuming retrieval is the only problem. If fixed input and reserves already exceed the window, removing every retrieved chunk is still insufficient.
- Treating room as relevance. Filling a large window does not establish that a model will find, trust or correctly use the right evidence.
From a plan to a real request
Find the current limits for the exact model and endpoint. Use its supported tokenizer or request-counting endpoint, including the structured messages, tools and any multimodal content. Counts from separate pieces may differ from the assembled request because formatting and boundaries matter.
Set a deliberate output allowance, verify how reasoning is counted, then choose a margin appropriate to your integration. Inspect actual usage and truncation/length outcomes after a request. Recount when the model, tool definitions, retrieved text or history policy changes.
For a multi-step agent, repeat the check at each model invocation as history and tool results grow. This single-request planner does not impose a runtime stop rule, prevent prompt injection, reserve a monetary budget or implement shared limits across workers.
Sources and next steps
Provider documentation checked October 8, 2026. The formulas and examples are original explanatory material; this page intentionally avoids model-specific presets that could silently become stale.
- OpenAI: managing the context window — distinguishes input, output and reasoning usage.
- Anthropic: token counting — documents supported request structures and notes that counts are estimates.