The Context Window Problem
Every LLM has a limit. GPT-4 Turbo gives you 128K tokens. Claude offers 200K. Sounds like plenty until you build a real application.
A 30-minute customer support conversation easily hits 50K tokens. Add retrieval context, system prompts, and tool definitions - you are at 80K before the user says hello.
This post documents the strategies we use to manage context effectively.
Understanding Context Economics
| Component | Typical Size | Priority |
|---|
| System prompt | 500-2000 tokens | Fixed |
| Tool definitions | 1000-5000 tokens | Fixed |
| Retrieved context | 2000-10000 tokens | Variable |
| Conversation history | Grows unbounded | Must manage |
| Current query | 50-500 tokens | Fixed |
| Response headroom | 1000-4000 tokens | Reserved |
For a 128K context window:
| Reserved | Purpose | Tokens |
|---|
| System + tools | Fixed overhead | 6,000 |
| Response | Generation space | 4,000 |
| Safety margin | Avoid truncation | 2,000 |
| Available | History + retrieval | 116,000 |
116K sounds like a lot. But a 2-hour support session generates 200K+ tokens of conversation.The simplest approach: keep recent messages, drop old ones.
Keep the last N messages or K tokens:
| Parameter | Trade-off |
|---|
| Window size | Larger = more context, higher cost |
| Unit | Messages vs tokens |
| Overlap | Keep some old context for continuity |
| Window Size | Context Quality | User Satisfaction |
|---|
| Last 5 messages | Poor | 2.8/5 |
| Last 20 messages | Moderate | 3.6/5 |
| Last 50 messages | Good | 4.1/5 |
| Token-based (32K) | Good | 4.2/5 |
| Scenario | Sliding Window Works? |
|---|
| Quick Q&A | Yes |
| Multi-turn tasks | Partially |
| Long sessions | No - loses important context |
| Reference-heavy | No - forgets references |
Compress old context instead of dropping it.
| Approach | Compression | Quality |
|---|
| LLM summary | 10:1 | High |
| Extractive | 5:1 | Medium |
| Key points only | 20:1 | Medium |
| Hierarchical | 50:1 | High |
Summarize as conversation grows:
| Trigger | Action |
|---|
| Every N messages | Summarize oldest chunk |
| Token threshold | Compress when approaching limit |
| Topic change | Summarize previous topic |
| Metric | Target |
|---|
| Information retention | Over 90% of key facts |
| Compression ratio | At least 5:1 |
| Coherence | Readable standalone |
| Latency overhead | Under 2 seconds |
| Approach | Context Preserved | User Satisfaction |
|---|
| No summarization | 100% (until dropped) | 3.2/5 |
| Basic summary | 75% | 3.9/5 |
| Hierarchical summary | 85% | 4.4/5 |
Different levels of detail for different time horizons.
| Tier | Content | Retention | Detail |
|---|
| Working | Current turn + recent | Full | Complete messages |
| Short-term | Last hour | Summarized | Key points |
| Long-term | Session history | Highly compressed | Facts and decisions |
| Persistent | Cross-session | Extracted | User preferences |
| Trigger | From | To | Action |
|---|
| 10 messages | Working | Short-term | Summarize chunk |
| 1 hour | Short-term | Long-term | Extract key facts |
| Session end | Long-term | Persistent | Store important info |
Context Assembly
When building the prompt:
| Priority | Source | Tokens |
|---|
| 1 | Working memory | 8,000 |
| 2 | Short-term summary | 2,000 |
| 3 | Long-term facts | 1,000 |
| 4 | Persistent context | 500 |
| 5 | Retrieved context | 4,000 |
Strategy 4: Retrieval-Augmented Context
Do not remember everything - retrieve what you need.
| Trigger | Retrieval |
|---|
| User mentions topic | Related conversation chunks |
| Tool call needed | Previous similar tool uses |
| Reference detected | Original referenced content |
| Question asked | Relevant past Q&A |
Index conversation chunks for retrieval:
| Field | Purpose |
|---|
| content | Searchable text |
| embedding | Semantic search |
| timestamp | Recency ranking |
| topic | Topic filtering |
| entities | Entity lookup |
Retrieval vs Full History
| Metric | Full History | Retrieval |
|---|
| Token usage | 50K+ | 5-10K |
| Relevance | Mixed | High |
| Cost | High | Low |
| Accuracy on references | 100% | 95% |
Guide the model to focus on what matters.
| Technique | How It Works |
|---|
| Importance markers | Tag key messages |
| Recency weighting | Emphasize recent |
| Topic headers | Group by topic |
| Summary prefixes | "Previously discussed:" |
Structure context for better attention:
| Section | Format |
|---|
| System context | Clear role definition |
| Persistent facts | Bulleted list |
| Recent summary | Narrative paragraph |
| Active context | Full messages |
| Current query | Clearly marked |
We combine multiple strategies.
| Component | Strategy | Purpose |
|---|
| Recent messages | Sliding window (20 msgs) | Full detail |
| Older messages | Rolling summary | Compressed context |
| Important facts | Long-term extraction | Persistent memory |
| On-demand | Retrieval | Reference lookup |
Context Budget Allocation
| Component | Tokens | Percentage |
|---|
| System + tools | 6,000 | 5% |
| Working memory | 24,000 | 19% |
| Short-term summary | 4,000 | 3% |
| Long-term facts | 2,000 | 2% |
| Retrieved context | 8,000 | 6% |
| Response headroom | 4,000 | 3% |
| Available for retrieval | 80,000 | 62% |
| Metric | Before | After |
|---|
| Max conversation length | 50 messages | Unlimited |
| Context relevance | 65% | 89% |
| Reference accuracy | 72% | 94% |
| User satisfaction | 3.4/5 | 4.5/5 |
| Token cost per session | High variance | Predictable |
| Strategy | Added Latency |
|---|
| Sliding window | 0ms |
| Basic summary | 500-1500ms |
| Hierarchical | 200-500ms (amortized) |
| Retrieval | 50-200ms |
| Approach | Quality | Cost | Latency |
|---|
| Full history | Highest | Highest | Lowest |
| Aggressive summarization | Medium | Lowest | Medium |
| Hybrid | High | Medium | Medium |
| Edge Case | Handling |
|---|
| User references old message | Retrieve from index |
| Topic circles back | Include topic summary |
| Contradictory statements | Keep both, note conflict |
| Very long single message | Truncate or summarize inline |
- 1Context is a budget - Plan your token allocation like you plan any resource.
- 2Sliding window is not enough - Works for short conversations, fails for long ones.
- 3Summarization preserves more than dropping - 85% retention vs 0% for dropped messages.
- 4Hierarchical memory scales - Different detail levels for different time horizons.
- 5Retrieval beats storage - Do not remember everything, retrieve what matters.
- 6Hybrid approaches win - No single strategy handles all cases. Combine them.
Context window management is not glamorous, but it is the difference between an AI that forgets mid-conversation and one that maintains coherent multi-hour sessions.