Today's Overview
- KV Compression Windows Create Periodic Retrieval Weaknesses: The same information can vary by up to 40 percentage points in accuracy across phases. Deployments need phase sweeps, not just average scores.
- Writable Context Shifts Information Control to the Model: Context Language Models raise BrowseComp-Plus accuracy by 11.4% while cutting compute by 21.5%.
- Follow-Up Questions Split Into Breadth, Depth, and Stopping: With the same retrieval budget, Q&D beats prompting and outperforms a 15-times-larger prompted model on two of three multi-hop QA benchmarks.
Featured
01 Efficiency: Chunked Compression Shows Phase Gaps of Up to 40 Points
The same fact can swing by up to 40 percentage points in long-context retrieval accuracy because of its phase inside a compression window. Chunked key-value cache compression merges consecutive tokens at a fixed stride. That quietly adds a new coordinate: the information's phase relative to the window boundary.
This is not ordinary compression loss. The weakness repeats by position. Aggregate scores can look healthy while specific phases fail systematically. Researchers trained multiple models from scratch with different KV compression schemes, and every design reproduced phase sensitivity.
The abstract's causal interventions show that separate attention components favor retrieval from different phases. Idealized models suggest training gradients may actively create this sharp specialization. Whether that mechanism generalizes still requires the full paper. Hold content fixed and sweep every window phase before deployment.
Key takeaways:
- Chunked KV cache compression introduces a positional reliability risk, not just the usual accuracy-efficiency tradeoff.
- Average accuracy can hide phase gaps as large as 40 percentage points.
- Test retrieval across every compression-window phase before deployment.
Source: Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
02 Architecture: When the Model Owns Context
The paper contrasts CLMs with context management under external harness control. Context Language Models turn context into a file the model can rewrite. Information selection and reorganization become model behaviors.
This control lets the model learn what is most important to maintain in context. The abstract reports an 11.4% accuracy gain on BrowseComp-Plus with 21.5% less compute. CLMs also extend naturally to multiple agent contexts stored as files.
Unrestricted file updates introduce an engineering question the abstract does not answer. Before deployment, test whether updates are observable and recoverable.
Key takeaways:
- Context management is shifting from external rules to a model capability.
- Model-managed context can improve accuracy while reducing FLOPs.
- The supplied abstract does not report observability or recovery tests for unrestricted context updates.
Source: Context Language Models
03 Agent: Knowing How Far to Ask
Horizontal proactivity fills gaps already visible in the current context. Vertical proactivity follows dependency chains to needs that only earlier evidence can reveal. A stopping policy decides when to quit asking.
The paper reconstructs a needs graph from benchmark decomposition. This supports direct evaluation from interaction logs, without a model judge. Q&D needs no reward model.
At the same retrieval budget, its trained questioner beats prompting with the same model on all three multi-hop QA benchmarks. On two, it also beats a prompted model 15 times larger. Without extra training, it transfers to simulated customer support and completes more tasks with fewer questions.
Key takeaways:
- Define current gaps, downstream needs, and explicit stopping conditions separately.
- Needs graphs and evidence coverage can evaluate follow-up questions without another model.
- At equal retrieval spend, a trained small questioner may outperform a prompted model 15 times larger.
Source: Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

Also Worth Noting
Today's Observation
Together, the three featured papers expose an information-control plane that aggregate scores can hide. Compression windows determine whether existing information remains retrievable. Model-managed context decides what stays available. Questioning policy decides whether missing evidence gets acquired.
Each control fails conditionally. Retrieval depends on phase, retention on model edits, and acquisition on need depth and stopping. A single task success rate cannot establish context-system reliability.
The next evaluation should keep a separate information-control ledger. Record retrieval rates by compression phase and evidence acquisition by need level. Since models can update context without restrictions, separately test whether those updates are observable and recoverable.