[Submitted on 29 Jun 2026 (v1), last revised 23 Jul 2026 (this version, v4)]
Abstract:Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Recent systems make context management agent- or system-controlled, but they either learn a compression policy that discards evidence or manage context in a layer the agent never sees. We argue both leave a more basic gap unaddressed. Frontier language models are proprioceptively blind to their own context. From the prompt alone they cannot see how large, how old, or how used each block is, the signals a keep-or-drop decision needs. We hypothesize that competent context management is already latent in capable models, and that what is missing is not a learned policy but an interface exposing this state. We introduce VISTA (Visible Internal State for Tool Agents), a training-free, model-agnostic layer that represents working memory as typed, addressable blocks, surfaces a runtime dashboard of per-block token usage, recency, and access history, and archives blocks as recoverable full-fidelity payloads. On LOCA-Bench, BrowseComp-Plus, and GAIA, the same untrained interface transfers across 1M-, 100K-, and 10K-scale trajectories. On LOCA-Bench it improves four backbones and lifts Gemini-3-Flash from 22.7 to 50.7%. The lift grows with context pressure and transfers across backbones. Ablations further confirm that the dashboard matters beyond archive and recovery tools.
Submission history
From: Binyan Xu [view email]
[v1]
Mon, 29 Jun 2026 09:13:38 UTC (1,407 KB)
[v2]
Sun, 5 Jul 2026 16:39:48 UTC (1,555 KB)
[v3]
Thu, 16 Jul 2026 07:52:40 UTC (1,655 KB)
[v4]
Thu, 23 Jul 2026 11:45:15 UTC (1,748 KB)
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.