Pivot-SD Wins With 200 Problems; Queen Gains 915 Elo

Today's Overview

  • Pivot-SD Beats Full-Sequence SFT With Sparse Supervision. It used 200 problems and four samples each to outperform full-sequence SFT and same-budget diffusion RL. Results cover only math and coding benchmarks on LLaDA-8B-Instruct.
  • Better Text Alignment Can Weaken Physical Regression. Contrastive training improved retrieval across 8,000 controlled simulations from 80 families, but hurt physical attribute regression. Better reference-video retrieval also improved MiniMax-H3’s physical fidelity.
  • Queen Climbs From 1782 to 2697 Elo in Seven Rounds. The 4B chess language model connects a chess encoder to an instruction-tuned language model through cross-attention. Language-model-based fluency and coherence scores cannot establish explanatory faithfulness.

Featured

01 Training: High-Information Tokens Beat Full Sequences

Pivot-SD beat full-sequence supervised fine-tuning and same-budget diffusion RL with only 200 problems and four samples per problem. Instead of supervising every token, it selects a few high-impact commitments.

Masked diffusion models face a credit-assignment problem. Some positions sharply reduce uncertainty across the remaining masked sequence, yet standard post-training treats most of the trajectory almost equally. Pivot-SD identifies these commitments through information gain.

Successful trajectories receive cross-entropy reinforcement at those positions. Failed trajectories get targeted unlikelihood training there, while their remaining positions receive no update. High information gain does not prove universal causal importance across models.

The abstract reports gains on LLaDA-8B-Instruct math and coding benchmarks. Transfer to other diffusion models, larger datasets, and bigger training runs still needs testing. Resource-constrained teams may gain more from better credit assignment than from simply increasing sample volume.

Key takeaways:

  • Prioritize commitments that sharply reduce uncertainty across the remaining sequence.
  • Apply different local signals to successful and failed trajectories instead of fitting both uniformly.
  • The small-data result is promising, but should not be extrapolated beyond LLaDA-8B-Instruct yet.

02 Video Gen: Better Text Alignment Can Weaken Physical Regression

Contrastive training on physics video-text pairs improved retrieval and pair classification, but physical attribute regression got worse. Semantic alignment and recoverable quantitative information are different capabilities.

The benchmark contains 8,000 controlled simulations across 80 families. Evaluated pretrained omni-modal embedding models performed poorly on retrieval. Within-family video-description matching hovered near chance, while lightweight probes still recovered useful physical data from frozen video representations.

Feeding retrieved reference videos into MiniMax-H3 improved the physical fidelity of generated videos. Stronger retrieval models produced larger gains in these experiments. The size of those gains, and their transfer to real footage, need broader testing.

Video teams should test semantic retrieval, physical recoverability, and final generation quality separately. One aggregate score is not enough.

Key takeaways:

  • Contrastive training may trade quantitative physical information for stronger text alignment.
  • Evaluate retrieval, physical attribute regression, and within-family video-description matching as separate capabilities.
  • Add reference-video retrieval to generation pipelines, but validate physical fidelity at the output.

03 Interpretability: Connect a Silent Expert to a Language Model

Queen is a 4B-parameter chess language model. Cross-attention connects a specialized chess encoder to an instruction-tuned language model, which extracts chess concepts from the encoder’s representations.

The system analyzes positions after candidate moves and folds those conclusions into its current explanation. It then distills the result back into the model. Seven iterations raised playing strength from 1782 to 2697 Elo.

That gain suggests a “silent expert plus language interface” may suit specialist domains better than a larger language model alone. Yet the abstract reports only language-model-based fluency and coherence evaluations.

Coherence approached GPT-5.6-Sol (high), but that does not prove the explanations reflect the chess encoder’s actual decision basis. Robotics and computer-use systems with specialist encoders should test this design alongside dedicated faithfulness checks.

Key takeaways:

  • When a reliable specialist model already exists, build a language interface before retraining a generalist.
  • A gain of more than 900 Elo shows that iterative distillation can convert future-position analysis into stronger decisions.
  • Evaluate linguistic coherence and decision faithfulness separately.
Pivot-SD Wins With 200 Problems; Queen Gains 915 Elo

Also Worth Noting

04
Static Reconstruction Does Not Transfer to Time-Varying Scenes Evaluation4DCodeBench asks agents to turn video deformation, fluids, and fractures into executable graphics programs. link
05
Target Frames Help Autoregressive Video Models Plan Ahead Video GenProAR constrains long-range outcomes and short-range transitions separately. It surpassed a fully trained standard autoregressive baseline using only 25% of the training steps. link
06
Fine-Grained Multi-Model Selection Beats Single-Model Bias Baselines SafetyCBM selects and organizes multiple models by fine-grained behavior. Its gains must be weighed against the inference cost of multi-model coordination. link
07
Action Policies Can Pretrain Without Action Labels RoboticsNAVA-WAM transfers supervision from future video changes into policy learning, reducing dependence on labeled robot trajectories. link
08
Program Evolution Should Optimize Improvement per Cost Code IntelligenceFrugalEvo assigns strategy exploration to stronger models and implementation to cheaper ones. BA-AUC explicitly measures progress across the cost curve. link
09
Browser Agents Now Need Cross-Modal Evidence Chains EvaluationHyperBrowseComp contains 423 human-written and verified questions across 13 languages. Sources include videos, scanned documents, images, maps, and obscure material. link
10
Ordinal Filtering Can Replace Repeated Group Sampling TrainingFTW reuses trajectories through a replay buffer. It trades CPU memory for fewer samples in critic-free reinforcement fine-tuning for stateful environments. link
11
Models Can Learn Which Modality to Acquire First MultimodalECHO-k uses pretrained internal representations as a proxy objective, producing acquisition orders independent of labels and specific downstream tasks. link
12
Reasoning Data Synthesis Changes Tasks and Tooling Together TrainingOnline improvement converts intermediate failures into reusable skills. Post-batch updates revise skills, prompts, and workflows while model weights and verification criteria stay fixed. link
13
Continually Adapting World Models Should Forget Stale Facts EvaluationThe paper separates permanent physical laws from revisable instance facts by invariance timescale. It proposes reporting both invariant regression and revision delay. link
14
Unlabeled Target Recordings Can Generate Multimodal Adapters MultimodalZeroMAG adds accompanying physiological signals while keeping the EEG encoder frozen. It trails supervised multimodal adaptation by only 0.50 percentage points on average. link
15
A Restricted Decoder Preserves More Transferable 3D Geometry MultimodalSNAP combines a local decoder with latent reconstruction. Under camera shifts, it degrades more slowly than standard 2D representations. link