Today's Overview
- Causal Pretraining Determines Whether Long Histories Improve Control: Extending video history from zero to 19.2 seconds raised RoboCasa GR-1 success from 63.3% to 78.7%. Bidirectional pretraining produced no net gain. Streaming execution kept action-chunk latency to 107.4 milliseconds on an RTX 5090.
- Visual Document Vectors Are Not Harmless Copies: Attackers recovered 47% of text and 45% of sensitive words. They could also use 98.4% of reconstructed pages to retrieve the source page. Shuffling can be reversed, while pooling has resisted attacks so far.
- Hybrid Attention Depends on How You Scale: Sliding-window hybrids suit training-free length extrapolation. Linear-attention hybrids benefit more from continued long-context pretraining. A 16× extrapolation result on 64k NIAH-SK1 does not prove broad capability.
Featured
01 Long Context Needs Causal Pretraining
With autoregressive pretraining, extending video history from zero to 19.2 seconds raised RoboCasa GR-1 success from 63.3% to 78.7%. A bidirectionally pretrained initialization produced no net gain. What matters is not how much history the model sees, but whether pretraining teaches how the past produces the future.
Long-WAM learns causal prediction from robot and egocentric videos without action labels. It then combines that structure with robot-domain pretraining and preserves it during action adaptation. The model achieved higher success rates across several long-horizon tasks.
Deployment also required streaming encoding, asynchronous execution, and hardware acceleration. These kept action-chunk latency to 107.4 milliseconds on an RTX 5090, including future video-latent prediction. Long-WAM reached 95% success in real-world moving-cup stacking. Pi0.5 and Fast-WAM each failed all 20 attempts. Broader tasks and hardware still need full-paper scrutiny and independent replication.
Key takeaways:
- Check whether performance improves with history length, not only the maximum supported window.
- Autoregressive video pretraining may suit control tasks that require progress and motion tracking better than bidirectional representations.
- Evaluate long-horizon robot systems on both success rate and end-to-end action latency.
Source: Long-WAM: Scaling the Context of World-Action Models
02 Embeddings Are Still Sensitive Data
A visual document page often becomes roughly 1,000 patch vectors arranged in raster order. That structure preserves layout and semantic clues, not just unreadable features. Researchers recovered 47% of the original text and 45% of sensitive words from raw indexes.
The reconstructions were far from lossless, yet 98.4% retrieved their source page at rank one. Pooling and order shuffling both reduced text recall to about 8%. The defenses did not hold equally well.
After recovering vector order, the attack raised source-page rank-one retrieval from 3.8% to 93.5%. Researchers have not yet reversed pooled indexes successfully. Without tuning, the attack also transferred to another multi-vector retriever. There, 70.2% of reconstructed pages identified the source, although text recovery trailed a nearest-neighbor baseline.
Key takeaways:
- Treat visual document indexes as sensitive data, not sanitized copies.
- Do not rely on vector-order shuffling. Attackers may reconstruct the sequence.
- Test pooling against both retrieval-quality loss and privacy attacks before treating it as a security boundary.
Source: Inverting Multi-Vector Visual Document Indices
03 Choose Attention Based on Your Scaling Plan
No hybrid-attention recipe covers every long-context goal. First decide whether you need immediate length extrapolation or stronger long-text ability through continued pretraining.
Hybrid sliding-window attention performed better at training-free length extrapolation. Hybrid linear attention gained more from continued long-context pretraining. Sliding-window models may also need larger windows during further training, since short windows can restrict long-range learning.
The authors also propose sliding-window linear attention. It achieved 16× training-free extrapolation with 100% accuracy on 64k NIAH-SK1. That is one needle-in-a-haystack test, not proof of stronger reasoning, retrieval, or generation.
Key takeaways:
- Prioritize sliding-window hybrids when you need context expansion without additional training.
- Focus on linear-attention hybrids when planning continued long-context pretraining.
- Reproduce the 16× result on realistic long-document tasks before drawing broader conclusions.
Source: Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position

Also Worth Noting
Today's Observation
Today's three featured papers carry the same warning: changing a representation or adding capacity does not guarantee the expected capability. Robot histories of 19.2 seconds improved control only when pretraining preserved causal prediction. Sliding-window and linear attention favor training-free extrapolation and continued pretraining, respectively. Visual document vectors may look unreadable, yet retain enough structure to recover sensitive text, layouts, and source pages.
Window length, vector format, and architectural complexity cannot replace validation. Split acceptance into two tracks. Task tests should measure what the new representation can still do. Threat tests should measure what an attacker can recover or infer. For every representation change, add one task ablation and one adversarial recovery test. Make both release gates.