Researchers from Meta, MIT, and the University of Washington introduce Context Language Models (CLMs), a new approach that enables language models to manage and edit their own context rather than relying on predefined mechanisms for summarization, compression, and information retrieval, reporting substantial gains in both performance and computational efficiency.
The researchers implemented the mechanism allowing LLMs to manage their own context by treating it as a file that they can update with no restrictions. This approach offers an alternative to conventional context-management strategies used to prevent context from growing indefinitely, such as compaction, offloading, and retrieval. Each of these techniques has its own limitations. Summarization can discard critical details or introduce inaccuracies, compaction often uses a predefined set of rules, and external memory requires agents to decide what to bring back into the context.
A CLM can instead rewrite old messages, preserve important facts, remove irrelevant information, maintain progress notes, and learn its own context-management strategies. The researchers emphasize that LLMs can learn strategies that go beyond existing human-designed approaches, potentially surpassing "existing human priors".
By shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable in-context learning and parametric learning for context management. We first show that users can steer context management simply by telling the agent their desired strategy. In addition, CLMs can evolve an in-context skill document that captures useful context-management procedures for future reuse.
New behaviors that CLM can learn include creating internal notes, removing irrelevant intermediate results while preserving useful ones, tracking unsuccessful experiments alongside ideas to explore further, and more.
The researchers explored three different approaches to context management. First, zero-shot context management, where the CLM received no specific training. Second, in-context learning, where the CLM was instructed using natural language on how to manage its context and refined its strategy through an iterative skill-optimization loop. Third, reinforcement learning, where the CLM learned context-management strategies using task success as the primary objective and computational efficiency as an additional optimization criterion.
The researchers evaluated CLMs across several benchmarks, reporting substantial gains in both performance and computational efficiency. Zero-shot CLMs achieved:
11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task
CLMs using in-context learning improved accuracy on ContextBench tasks by up to 35.9% points at lower compute. Finally, reinforcement learning improved Qwen3.5-9B performance on BrowseComp-Plus from 28.8% to 42.5% with a 47.6% relative improvement while using 12% fewer FLOPs.
The researchers acknowledge several unresolved challenges. First, a CLM may discard important information than cannot be later retrieved. Second, allowing CLMs to edit their own context introduces new safety risks, as the editable context "can become another channel through which prompt injections or self-generated instructions persist across turns". Most importantly, greater control over context does not necessarily translate into better behavior, as models may still make poor decisions about what to retain, modify, or discard.
Commenting the announcement on X.com, @Rennix7t warns that "the accuracy and computational power results in the paper come from specified tasks", while @omarsar0 describes the approach as an interesting research direction but expresses reservations about trusting a model to manage its context end-to-end, arguing that better solutions are still needed.
On Reddit, Combinatorilliance argues that CLMs are not yet production-ready highlighting that rewriting the context causes the model cache to be invalidated and that the "caching trick [used by the researchers, EN] isn't ideal either and has some drawbacks that need to be taken into consideration".