All Articles
CategoryAI
Reading Time
15 min read
Published
2026-10-07
Word Count
3,706words

Grab a coffee — this one is a deep dive!

Context Management in Claude: Compaction or Memory Tool

Summary

Context management in long agent sessions: clear_tool_uses_20250919 defaults, carrying persistent information with the memory tool, and a comparison of compaction strategies.

  • clear_tool_uses_20250919 triggers at a default 100,000 input token threshold and preserves the 3 most recent tool use/result pairs.
  • clear_thinking_20251015 must be listed first in the edits array; its default behavior varies by model class (Opus 4.5+/Sonnet 4.6+/Fable/Mythos differ).
  • The memory tool (memory_20250818) runs client-side and, combined with context editing, can rescue information to the /memories directory before it's deleted.
  • The old compaction_control was removed from the Python SDK in v1.0.0 (August 20, 2026) and deprecated in TS/Ruby, replaced by the tool runner's compact_before_next_turn()/compactBeforeNextTurn() methods and server-side on-demand compaction.
Context Management in Claude: Compaction or Memory Tool

A long-running Claude agent fills its context window a little more with every tool call, and the growth usually goes unnoticed: file reads, search results, and intermediate thinking steps pile up, the prompt cache prefix breaks, and the same task starts costing more tokens than the last run. Anthropic doesn't solve this with a single tool but with three complementary mechanisms: server-side context editing, the memory tool, and SDK compaction (which now lives on in a different form). Under the umbrella of LLM context management, this article clarifies where compaction and the memory tool diverge, which mechanism kicks in when, what their defaults are, and how to choose between them.

💡 Pro Tip: Before adding the clear_tool_uses_20250919 strategy, put any tools whose results you still reference (e.g. a file-read result) into the exclude_tools list — otherwise Claude may keep referring to a result it produced but can no longer see.

Table of Contents

Why context quietly gets expensive

An agent session isn't a single message exchange — it's a chain of accumulating tool calls. Every tool result, thinking block, and intermediate response gets appended to the context window, and as the conversation grows this buildup produces two costs: raw token count rises, and the prompt cache prefix must be carried forward with every addition. Anthropic's official documentation describes server-side context management as the primary strategy for long-running conversations, since the problem should be solved by a mechanism the API itself understands, not hand-written summarization code on the client.

The critical distinction: "shrinking context" isn't one operation — it's at least three operations serving three different purposes. Deleting a tool result, deleting a thinking block, and summarizing the whole conversation differ sharply in loss tolerance — one discards a reproducible result, the other performs an irreversible summary. Saying "I'm using compaction" without seeing this difference means not knowing which information is permanently lost.

None of these three mechanisms is sufficient alone, because each optimizes a different cost axis. Tool result clearing lowers token count but breaks the prompt cache; thinking clearing lightens context but varies by model class; the memory tool adds no token cost but needs your application to have file-system access. Balancing these three axes when designing an agent architecture means picking a combination that fits the task's own loss tolerance, not finding one "best" setting.

Three strategies: context editing, memory tool, compaction

Anthropic's context editing documentation splits context management into three separate building blocks: tool result clearing, thinking block clearing, and client-side SDK compaction. The first two run server-side under the context_management parameter and are enabled with the context-management-2025-06-27 beta header; the third lives in the SDK's tool runner layer. On top of these, there's a completely separate layer, the memory tool (memory_20250818) — this tool does the _opposite_ of context editing: it moves information that's about to be deleted from context into persistent files before it disappears. Because context editing splits into two strategies of its own, the rest of this article counts these mechanisms as four separate items.

Mechanism
Where it runs
What it does
Reversibility
clear_tool_uses_20250919
Server-side
Replaces old tool results with a placeholder
Result is lost, the tool call stays visible
clear_thinking_20251015
Server-side
Clears old thinking blocks
Varies by model class
Memory tool (memory_20250818)
Client-side file system
Writes information to the /memories directory
Persistent, survives a reset
SDK compaction (compaction_control, tool runner)
Client-side (deprecated in TS/Ruby, removed in Python v1.0)
Replaces the conversation with a server-written summary
Detail outside the summary is lost

The right way to think about all three together is to ask which information is "cheap to reproduce": a file-read result can be requested again, so it's a good candidate for tool result clearing; but a user's constraint, or the reasoning behind a decision, can't be reproduced — that needs to be made persistent with the memory tool.

clear_tool_uses_20250919: settings and defaults

The clear_tool_uses_20250919 strategy kicks in once the conversation reaches a certain input token threshold, replacing the oldest tool use/result pairs with placeholder content. Per Anthropic's documentation table, the default trigger value is 100,000 input tokens, customizable in input_tokens or tool_uses. The default keep parameter is 3 tool use/result pairs — the 3 most recently used pairs always stay in context, the rest get cleared.

Four parameters make a practical difference:

  • trigger: determines when the strategy activates (default 100K input tokens).
  • keep: determines how many of the most recent tool pairs are preserved (default 3).
  • clear_at_least: if the API can't clear at least this many tokens, the strategy isn't applied at all — meaning there's no "partial clearing."
  • exclude_tools: the list of tool names whose use/result pairs are never cleared.

The default behavior only clears tool _results_; as long as clear_tool_inputs is false, Claude's original tool calls (which tool, which parameters) stay visible. This lets the model remember "why it called this tool" while removing the result itself.

json
1{
2 "context_management": {
3 "edits": [
4 {
5 "type": "clear_tool_uses_20250919",
6 "trigger": { "type": "input_tokens", "value": 100000 },
7 "keep": { "type": "tool_uses", "value": 3 },
8 "clear_at_least": { "type": "input_tokens", "value": 5000 },
9 "exclude_tools": ["memory"]
10 }
11 ]
12 }
13}

The clear_at_least field here matters especially for agents using prompt caching: the clearing operation invalidates the cached prefix, so instead of "clearing a small amount and breaking the cache for nothing," either a sufficiently large block gets cleared or nothing gets cleared at all.

clear_thinking_20251015 and its effect on prompt cache

Clearing thinking blocks is a separate strategy because their default behavior varies by model class. Per Anthropic's documentation, Opus 4.5+ and Sonnet 4.6+ models preserve all previous thinking blocks; Fable and Mythos preserve all turns; older model classes keep only the last turn. So "does my context automatically bloat with extended thinking" depends on the model version.

Ordering matters when both strategies are used together: the documentation states that clear_thinking_20251015 must be listed first in the edits array when multiple strategies are used; it doesn't say why, so apply the rule as given.

The distinction that matters for prompt cache is this: in Claude Fable 5.1 and Claude Opus 5.5, server-side context management doesn't invalidate thinking blocks. By contrast, manual edits made to previous turns on the client side can invalidate the thinking blocks in every subsequent assistant turn. In other words, clear_thinking_20251015 is designed to be cache-friendly, while manually editing past messages yourself carries the risk of breaking the cache.

python
1response = client.beta.messages.create(
2 model="claude-fable-5-1",
3 max_tokens=4096,
4 betas=["context-management-2025-06-27"],
5 context_management={
6 "edits": [
7 {"type": "clear_thinking_20251015"},
8 {
9 "type": "clear_tool_uses_20250919",
10 "trigger": {"type": "input_tokens", "value": 100000},
11 },
12 ]
13 },
14 messages=conversation_history,
15 tools=tool_definitions,
16)

Carrying persistent information with the memory tool

The memory tool (memory_20250818) works in exactly the opposite direction from context editing: it makes important information persistent before it gets deleted. According to Anthropic's memory tool documentation, the tool runs entirely client-side — Claude generates file operation requests (read, write, list), and it's your application that actually executes them; Anthropic stores no files server-side. This tool type is available on all Claude 4-and-later models.

This is where the real benefit shows up when context editing and the memory tool are used together: as the conversation approaches the tool-result-clearing threshold, Claude automatically gets a warning and can write important information to memory files before it's deleted. The two mechanisms aren't competitors — they're a pair that works in sequence, one lightening context while the other rescues information about to disappear.

json
1{
2 "tools": [{ "type": "memory_20250818", "name": "memory" }],
3 "context_management": {
4 "edits": [
5 {
6 "type": "clear_tool_uses_20250919",
7 "trigger": { "type": "input_tokens", "value": 100000 }
8 }
9 ]
10 }
11}

In practice this means a code review agent can write "the three critical findings in this file" into /memories/findings.md, then let the tool results from reading those files get cleared from context — the findings aren't lost, only the raw file contents are.

The memory tool being client-side is the most important architectural difference from context editing. Context editing runs entirely on Anthropic's server and needs no extra code from your application; the memory tool, on the other hand, needs a layer that actually executes the file operations — typically a local file system, an object store, or a database. That makes the memory tool a mechanism requiring more engineering decisions than context editing, but in return it can survive even across conversation sessions.

The fate of client-side compaction: removal or redesign

Two separate facts need to be kept apart here. Anthropic's context editing documentation states that the old compaction_control parameter is deprecated in the TypeScript and Ruby SDKs and will be removed in the future, printing a deprecation warning when enabled. In the Python SDK the situation is more definitive: the compaction_control argument and CompactionControl type were completely removed in v1.0.0, released August 20, 2026 — per the Python SDK's migration guide, this change moves summarization to server-side compaction inside the API itself, instead of an extra client round-trip.

But this doesn't mean "compaction is gone from the SDKs entirely." The Python SDK's changelog shows that between September 15–22, 2026 (just weeks before this article), a new compaction mechanism arrived in the SDK: v1.6.0 added signed compaction blocks and the beta compaction parameter on the API side; the tool runner method came in v1.7.0, which added compact_before_next_turn() and refined its state and error handling; v1.8.0 fixed the compaction request's body. TypeScript saw a parallel change: v0.127.0 (September 18, 2026) added the equivalent compactBeforeNextTurn() method to the tool runner.

So the correct framing is: the old, fully client-side summarizing compaction_control mechanism is being deprecated in TS/Ruby and has already been removed in Python — but in its place, a new, thinner helper method has arrived on the tool runner that triggers server-side compaction. Saying "compaction is over" is wrong; "the old compaction mechanism is handing off to the server-side one" is a more accurate description. This is also the official migration path: to use server-side compaction with the tool runner, pass the compact_20260112 edit into the request's context_management parameter (the Python SDK migration guide's "After" example shows this edit with betas=["compact-2026-01-12"]) — and per the tool runner docs' warning, on a given runner use either the compact_before_next_turn() helper or the context_management compaction edit, never both. Server-side compaction comes in two forms: threshold compaction (compact_20260112, inside context_management) and on-demand compaction (compact-2026-09-04, the top-level compaction parameter); the docs say the two can't be combined on one request.

SDK
Old mechanism
Status
New mechanism
Python
compaction_control kwarg, CompactionControl type
Removed in v1.0.0 (Aug 20, 2026)
compact_before_next_turn() (tool runner, v1.7.0+)
TypeScript
compaction_control parameter
Deprecated, to be removed (prints a warning)
compactBeforeNextTurn() (tool runner, v0.127.0+)
Ruby
compaction_control parameter
Deprecated, to be removed (prints a warning)
No compact_before_next_turn helper in the Ruby tool runner; the server-side context_management compaction edit is used instead

Server-side on-demand compaction

The actual API feature underlying this redesign is on-demand compaction, added to the Messages API on September 14, 2026. Per Anthropic's API release notes, it's enabled with the beta compact-2026-09-04 header: it compacts a conversation on request and returns a signed compaction block, which replaces the compacted messages in later requests. Compaction replaces a conversation's old turns with a summary Claude writes on the server — the summarizing logic isn't a client-side algorithm, but a summary the model itself produces.

This explains why compact_before_next_turn() / compactBeforeNextTurn() exist in the tool runner: these methods schedule the next compaction request once the current turn finishes — the SDK no longer writes its own summary, it just bridges to trigger on-demand compaction on the server.

This distinction clarifies what changed versus the old architecture. In the old compaction_control approach, summarizing logic lived on the client: the SDK made a separate API call to summarize the conversation, formatted the result, and appended it to the next request — an extra round-trip plus summarizing code to maintain. With on-demand compaction, this moves entirely to the server: once the request carries the compact-2026-09-04 beta header, summarizing happens inside the API call, and the resulting signed compaction block can be used directly in later requests. The new tool-runner methods are simply a convenience layer that sends the compaction request and swaps the history for you, while you still decide when to compact.

Decision tree: which strategy, when

The three mechanisms aren't mutually exclusive; the question you need to ask is "which information's loss can I tolerate."

  • Is the tool result reproducible (a file read, a search result, an API call)? → Clear it with clear_tool_uses_20250919, and define exceptions with exclude_tools where needed.
  • Was the thinking process only needed in the moment, or does it influence later decisions? → Use clear_thinking_20251015 first; check the default keep behavior for your model class.
  • Is the information irreproducible and does it need to survive a conversation reset (a user decision, a project constraint, a past error analysis)? → Write it to the /memories directory with the memory tool.
  • Has the whole conversation grown too long, and clearing individual tool results isn't enough? → Trigger on-demand compaction with the compact-2026-09-04 beta header; if you're on an SDK, use compact_before_next_turn() / compactBeforeNextTurn(), and don't rely on the old compaction_control.
  • Is the prompt cache hit rate critical? → Tool result clearing breaks the cache prefix; use clear_at_least to prefer larger, less frequent clears over small, frequent ones.

Answering these five questions in order turns "which setting should I turn on" into "which information am I willing to lose" — which is the real decision anyway.

One more practical consequence of this ordering: turning on all four strategies at once isn't the same as turning them on without thinking about each individually. Use clear_tool_uses_20250919 and clear_thinking_20251015 together, and the edits order is mandatory; use the memory tool with context editing, and you need to review exclude_tools; enable on-demand compaction, and you need to check your tool runner version (v1.7.0+ in Python, v0.127.0+ in TypeScript). So the decision tree answers not just "which strategy," but also "in what order and with what version requirements."

A practical scenario: configuration for a file-editing agent

Consider a file-editing agent: its file-read results are large and reproducible, but a user constraint like "don't touch this file" needs to be permanent. The configuration below brings the three strategies from earlier sections together into a single request:

typescript
1const response = await client.beta.messages.create({
2 model: "claude-opus-5",
3 max_tokens: 4096,
4 betas: ["context-management-2025-06-27"],
5 tools: [
6 { type: "memory_20250818", name: "memory" },
7 { type: "text_editor_20250728", name: "str_replace_based_edit_tool" },
8 ],
9 context_management: {
10 edits: [
11 { type: "clear_thinking_20251015" },
12 {
13 type: "clear_tool_uses_20250919",
14 trigger: { type: "input_tokens", value: 100000 },
15 keep: { type: "tool_uses", value: 3 },
16 clear_at_least: { type: "input_tokens", value: 5000 },
17 exclude_tools: ["memory"],
18 },
19 ],
20 },
21 messages: conversationHistory,
22});

The exclude_tools: ["memory"] line is the key point here: the memory tool's own read/write results aren't deleted by tool result clearing — because memory is already the persistence layer, and clearing it too would defeat the purpose.

clear_thinking_20251015 sitting first isn't a preference, it's a requirement: the documentation mandates this strategy be listed first in edits when multiple strategies are used, without stating why. Leave this line's position unchanged; the ordering stays the same with an editing tool like str_replace_based_edit_tool too.

GOLDEN TIP

The most valuable insight in this article

This tip holds the article's most important takeaway.

Easter Egg

You found a hidden gem!

There's a hidden detail in this section. Want to uncover it?

Reader Reward

Below we've collected the checklist items you need to go through when adding the four strategies covered in this article (tool result clearing, thinking clearing, memory tool, on-demand compaction) to an agent architecture; you can work through them one by one, checking each off in your own project.

FAQ

How is context managed in long agent sessions?

Anthropic offers three complementary mechanisms: server-side context editing (automatically clears old tool results and thinking blocks), the memory tool (stores information as a persistent file in the /memories directory, surviving even a conversation reset), and on-demand compaction (replaces a conversation's old turns with a server-written summary). All three can be used together; the choice depends on which information is reproducible and which needs to be persistent.

Should I use server-side compaction or tool-result clearing?

These are two separate layers, not a single choice. Tool result clearing (clear_tool_uses_20250919) replaces old tool results with a placeholder while preserving conversation flow and tool calls; on-demand compaction replaces old turns with a server-written summary — you can keep recent turns verbatim if you want — so it's more aggressive, uses fewer tokens, but risks more detail loss. Anthropic's documentation positions server-side strategies as generally preferred for long conversations.

What token threshold and how many calls does clear_tool_uses preserve?

The default trigger threshold is 100,000 input tokens (trigger can be customized in terms of input_tokens or tool_uses); by default, the 3 most recent tool use/result pairs are preserved (keep). The clear_at_least parameter guarantees a minimum amount of tokens to be cleared — if that amount can't be met, the strategy isn't applied at all. Specific tools can be kept out of clearing with exclude_tools.

How does the memory tool work together with context clearing?

When used together, as the conversation approaches the tool-result-clearing threshold, Claude automatically gets a warning and can write important information to memory files before it's deleted. So even when raw tool results get cleared from context, the critical information extracted from them stays accessible via the memory tool — the two mechanisms are a complementary pair.

Can I keep using the old compaction_control parameter?

Not in the Python SDK: with v1.0.0 (August 20, 2026), the compaction_control argument and CompactionControl type were completely removed. In TypeScript and Ruby the parameter still works but is marked deprecated and will be removed, printing a warning when enabled. Either way, migrating new projects to the tool runner's compact_before_next_turn() / compactBeforeNextTurn() methods keeps you compatible with server-side on-demand compaction.

Conclusion

Context management isn't a single "compress" button anymore — it's a set of layers where four separate tools (tool result clearing, thinking clearing, memory tool, on-demand compaction) are used together. The caching logic we covered in Claude Prompt Caching explains why the clear_at_least parameter exists here; for scenarios that analyze large codebases in one pass, take a look at Claude 1M Context Window — that piece covers the one-time large-read scenario, while this one covers the cost/quality axis of long-running agent sessions. For extended thinking's relationship to context, see our Claude Extended Thinking article, and for the general framework of the memory concept, our Claude Projects and Memory article. If you want to keep up with model version updates, our Claude Fable 5.1: What's New and What Breaks article can help you stay current.

In practice, the most effective approach is to see the four mechanisms not as competitors but as layers with different loss tolerances: cheaply clear what's reproducible, write what needs to be permanent to memory, and let the server summarize the rest.

Sources

Tags

#Claude#context editing#compaction#memory tool#prompt caching#agent architecture#Anthropic API
Muhittin Çamdalı

Muhittin Çamdalı

Lead Mobile Engineer

Lead Mobile Engineer with 12+ years of experience. Expert in iOS, Android and cross-platform architectures with Swift, SwiftUI, Kotlin and Flutter. I build performant, user-friendly mobile apps.

iOS Development News

Weekly Swift tips, SwiftUI tricks and iOS best practices. No spam, only valuable content.

Your subscription starts when you open the link in the confirmation email and press “Confirm my subscription”. The newsletter keeps open/click statistics; you can unsubscribe anytime with one click. Privacy

Share