LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via State Proprioception
Binyan Xu, Haitao Li, Kehuan Zhang
Why It Matters
What makes this one worth your time
Efficient context management is crucial for long-horizon tasks in language models, and VISTA offers a novel approach to improve this capability without additional training.
VISTA enables language model agents to self-manage context by providing a visible internal state.
Summary
The paper introduces VISTA, a model-agnostic layer that enhances language model agents by providing them with a visible internal state for better context management. This approach allows agents to make informed decisions about context retention and archiving, improving performance on several benchmarks.
Key contributions
- Introduction of VISTA, a training-free, model-agnostic layer for context management.
- Demonstration of VISTA's effectiveness across multiple benchmarks and scales.
- Provision of a runtime dashboard for token usage and context management.
Notable insights
- The concept of 'state proprioception' for language models, allowing them to be aware of their own context usage.
- A runtime dashboard that provides real-time insights into token usage and context status.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2606.30005v5 Announce Type: replace Abstract: Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Recent systems make context management agent- or system-controlled, but they either learn compression policies that discard evidence or manage context in a layer the agent never sees. We argue that both miss a more basic gap: frontier language models are proprioceptively blind to their own context. From the prompt alone they cannot reliably infer block size, recency, or the remaining budget, all of which are needed for keep-or-archive decisions. We introduce VISTA (Visible Internal State for Tool Agents), a training-free, model-agnostic layer that represents working memory as typed addressable blocks, surfaces a runtime dashboard of token usage, recency, archive status, and remaining budget, and archives blocks as recoverable full-fidelity payloads. On LOCA-Bench, BrowseComp-Plus, and GAIA, the same untrained interface transfers across 1M-, 100K-, and 10K-scale trajectories. On LOCA-Bench it lifts Gemini-3-Flash from 22.7 to 50.7%, reaches 58.0% on BrowseComp-Plus, and remains competitive on GAIA. Gains grow with context pressure and transfer across backbones, while ablations confirm that the dashboard matters beyond archive and recovery tools.