Free engineering tool · No sign-in or model calls

LLM context window calculator

A large language model (LLM) needs room for more than your question. Plan conversation history, tool data, retrieved evidence and the answer before your request runs out of context.

For retrieval-augmented generation (RAG) and agent workflows. Enter token counts, not words or characters. This is a budget planner, not a tokenizer, quality benchmark or cost estimate.

Plan one request

All counts below are in tokens, except the chunk count. Examples are illustrative workloads, not specifications for a named model.

The total context limit for your exact model and endpoint.

Instructions, examples and their formatting overhead.

The prior turns actually included in this request.

The new request, including any counted attachments.

Tool schemas and returned data included in the request.

Include reasoning here if it consumes the same output allowance. Do not count it twice.

Extra headroom you choose; it is not a provider guarantee.

A conservative per-chunk bound including source labels and separators.

The number of chunks you intend to include.

Planning result, not an API guarantee

1,600 tokens of headroom

14,400 planned / 16,000 context limit

The unfilled area is headroom after the output reserve and safety margin.
Fixed input
4,000
Retrieved content
6,400
Output reserve
3,000
Safety margin
1,000

Retrieved-chunk capacity

Maximum under these assumptions: 10 chunks. Requested: 8.

Capacity is not a recommendation to fill the window. More retrieved text can add distraction, duplication and latency.

The budget equation

For a model whose context allowance includes the input and generated tokens, reserve the answer before packing in retrieval:

fixed input = instructions + history + current message + tool data

available for retrieval = context window - fixed input - output reserve - margin

maximum chunks = max(0, floor(available for retrieval / tokens per chunk))

The output reserve includes reasoning when your endpoint counts it against that allowance. Confirm the exact model's separate input/output caps and reasoning rules. A request can satisfy this arithmetic and still violate a provider-specific limit.

Worked example: why ten chunks fit but eleven do not

Assume a 16,000-token window. These are teaching numbers, not a claim about a particular provider or model:

AllocationTokens
System instructions800
Conversation history2,500
Current message200
Tool definitions and results500
Output + reasoning reserve3,000
Safety margin1,000
Available for retrieved content8,000

At 800 tokens per chunk, 10 chunks fit exactly alongside the selected reserves. Eleven chunks require 8,800 tokens, exceeding the total plan by 800. Eight chunks leave 1,600 tokens of additional headroom. Try all three values in the calculator.

If chunk sizes vary, use a conservative per-chunk bound for this planning model. For precise packing, count each selected chunk with its labels and separators, assemble the complete request and count again. An average of 800 does not ensure that every group of ten will fit.

Five mistakes this arithmetic makes visible

  1. Spending the entire window on input. A large prompt still needs room for the response. Output limits and context limits are different constraints.
  2. Forgetting tools. Schemas, tool outputs and message wrappers can take substantial space. Tool use is not context-free.
  3. Counting words instead of model tokens. Tokenization varies by model, language and content. Do not treat a fixed characters-per-token rule as exact.
  4. Assuming retrieval is the only problem. If fixed input and reserves already exceed the window, removing every retrieved chunk is still insufficient.
  5. Treating room as relevance. Filling a large window does not establish that a model will find, trust or correctly use the right evidence.

From a plan to a real request

Find the current limits for the exact model and endpoint. Use its supported tokenizer or request-counting endpoint, including the structured messages, tools and any multimodal content. Counts from separate pieces may differ from the assembled request because formatting and boundaries matter.

Set a deliberate output allowance, verify how reasoning is counted, then choose a margin appropriate to your integration. Inspect actual usage and truncation/length outcomes after a request. Recount when the model, tool definitions, retrieved text or history policy changes.

For a multi-step agent, repeat the check at each model invocation as history and tool results grow. This single-request planner does not impose a runtime stop rule, prevent prompt injection, reserve a monetary budget or implement shared limits across workers.

Sources and next steps

Provider documentation checked October 8, 2026. The formulas and examples are original explanatory material; this page intentionally avoids model-specific presets that could silently become stale.