Causal Pretraining Makes Long Robot Context Pay Off

Today's Overview

  • Causal Pretraining Determines Whether Long Histories Improve Control: Extending video history from zero to 19.2 seconds raised RoboCasa GR-1 success from 63.3% to 78.7%. Bidirectional pretraining produced no net gain. Streaming execution kept action-chunk latency to 107.4 milliseconds on an RTX 5090.
  • Visual Document Vectors Are Not Harmless Copies: Attackers recovered 47% of text and 45% of sensitive words. They could also use 98.4% of reconstructed pages to retrieve the source page. Shuffling can be reversed, while pooling has resisted attacks so far.
  • Hybrid Attention Depends on How You Scale: Sliding-window hybrids suit training-free length extrapolation. Linear-attention hybrids benefit more from continued long-context pretraining. A 16× extrapolation result on 64k NIAH-SK1 does not prove broad capability.

Featured

01 Long Context Needs Causal Pretraining

With autoregressive pretraining, extending video history from zero to 19.2 seconds raised RoboCasa GR-1 success from 63.3% to 78.7%. A bidirectionally pretrained initialization produced no net gain. What matters is not how much history the model sees, but whether pretraining teaches how the past produces the future.

Long-WAM learns causal prediction from robot and egocentric videos without action labels. It then combines that structure with robot-domain pretraining and preserves it during action adaptation. The model achieved higher success rates across several long-horizon tasks.

Deployment also required streaming encoding, asynchronous execution, and hardware acceleration. These kept action-chunk latency to 107.4 milliseconds on an RTX 5090, including future video-latent prediction. Long-WAM reached 95% success in real-world moving-cup stacking. Pi0.5 and Fast-WAM each failed all 20 attempts. Broader tasks and hardware still need full-paper scrutiny and independent replication.

Key takeaways:

  • Check whether performance improves with history length, not only the maximum supported window.
  • Autoregressive video pretraining may suit control tasks that require progress and motion tracking better than bidirectional representations.
  • Evaluate long-horizon robot systems on both success rate and end-to-end action latency.

02 Embeddings Are Still Sensitive Data

A visual document page often becomes roughly 1,000 patch vectors arranged in raster order. That structure preserves layout and semantic clues, not just unreadable features. Researchers recovered 47% of the original text and 45% of sensitive words from raw indexes.

The reconstructions were far from lossless, yet 98.4% retrieved their source page at rank one. Pooling and order shuffling both reduced text recall to about 8%. The defenses did not hold equally well.

After recovering vector order, the attack raised source-page rank-one retrieval from 3.8% to 93.5%. Researchers have not yet reversed pooled indexes successfully. Without tuning, the attack also transferred to another multi-vector retriever. There, 70.2% of reconstructed pages identified the source, although text recovery trailed a nearest-neighbor baseline.

Key takeaways:

  • Treat visual document indexes as sensitive data, not sanitized copies.
  • Do not rely on vector-order shuffling. Attackers may reconstruct the sequence.
  • Test pooling against both retrieval-quality loss and privacy attacks before treating it as a security boundary.

03 Choose Attention Based on Your Scaling Plan

No hybrid-attention recipe covers every long-context goal. First decide whether you need immediate length extrapolation or stronger long-text ability through continued pretraining.

Hybrid sliding-window attention performed better at training-free length extrapolation. Hybrid linear attention gained more from continued long-context pretraining. Sliding-window models may also need larger windows during further training, since short windows can restrict long-range learning.

The authors also propose sliding-window linear attention. It achieved 16× training-free extrapolation with 100% accuracy on 64k NIAH-SK1. That is one needle-in-a-haystack test, not proof of stronger reasoning, retrieval, or generation.

Key takeaways:

  • Prioritize sliding-window hybrids when you need context expansion without additional training.
  • Focus on linear-attention hybrids when planning continued long-context pretraining.
  • Reproduce the 16× result on realistic long-document tasks before drawing broader conclusions.
Causal Pretraining Makes Long Robot Context Pay Off

Also Worth Noting

04
Keep the Existing Video DiT and Make Compressed Latents Compatible: Video GenGRACE adapts Wan2.1-I2V-14B at 480×832×81 resolution. It cuts tokens by 8× and latency by 11.1× while preserving the original pipeline's VBench quality. link
05
Separate Conflicting Gradients to Generate for 24 Hours: Video GenSGF+ assigns future-context writing and current-frame denoising to different parameters. Five-second training clips extrapolate to continuous videos up to 24 hours without extra data or long-video fine-tuning. link
06
Turn a Video Diffusion Model Into a Hand Tracker: RoboticsRLHND uses Cosmos 3 as a deterministic clip-level feature extractor. It estimates pose, contact, and force from monocular egocentric video, with real-robot experiments validating its use for robot learning. link
07
Text Rendering Benchmarks Now Test Whole Layouts: EvaluationUltraText Bench contains 432 bilingual Chinese-English prompts with four to 12 text regions. In tested settings, improved clarity sometimes reduced textual fidelity. link
08
Skill Libraries Need Retirement Policies: AgentSkillForge moves skills through trial, stabilization, and retirement so outdated experience stops misleading agents. It delivered relative gains of up to 7.8% across interactive benchmarks while controlling library size. link
09
Tool Use Maps to an Internal Direction: InterpretabilityAn inhibition-shaped mechanism compresses an agent's tool-use decision into one internal direction. It appears across seven Qwen, Mistral, and Granite models, offering an intervention point when agents discuss actions instead of executing them. link
10
Visual Token Pruning Can Strengthen Malicious Anchors: SafetyPruning may concentrate attention on hostile foreground content. The inference-time SAP defense reduced attack success by up to 62% in tested settings without sacrificing efficiency or utility. link
11
Reward Models Can Judge Without Writing Full Critiques: EfficiencyLatentGRM-8B compresses evaluation into continuous latent trajectories, shortening the judging process by about 9×. At vote@5, it cut total inference time by 6.1× to 7× while retaining competitive preference accuracy. link
12
Global Safety Scores Can Hide Local Regressions: EvaluationPatchBench-Local tests both harmful-neighbor repair and benign-neighbor preservation. It exposes severe local failures even when overall capabilities barely change. link
13
Neural PDE Surrogates May Learn Solver Errors: AI for ScienceFourier-symbol diagnostics show that surrogates can imitate training-solver discretization errors beyond 99.8% of the theoretical maximum for perfect imitation. link

Today's Observation

Today's three featured papers carry the same warning: changing a representation or adding capacity does not guarantee the expected capability. Robot histories of 19.2 seconds improved control only when pretraining preserved causal prediction. Sliding-window and linear attention favor training-free extrapolation and continued pretraining, respectively. Visual document vectors may look unreadable, yet retain enough structure to recover sensitive text, layouts, and source pages.

Window length, vector format, and architectural complexity cannot replace validation. Split acceptance into two tracks. Task tests should measure what the new representation can still do. Threat tests should measure what an attacker can recover or infer. For every representation change, add one task ablation and one adversarial recovery test. Make both release gates.