Today's Overview
- Relevant Documents Are Not a Good Evidence Set. AdaTutoRank adjusts tutoring intensity to candidate-set quality and distills guidance gains into per-token advantages, separating key evidence from “free-rider” documents.
- MALA Scores Every Valid Attention Pair Before Allocating Compute. In 128K-context operator tests, it accelerated training forward passes by 2.2×, backward passes by 3.0×, and decoding by 1.6×.
- Extreme Visual Compression Must Correct the Student’s Own Trajectory. With only 5% of visual tokens, LT-OPD raised Qwen3.5-4B’s average retained performance across nine benchmarks from 68.6% to 82.3%.
Featured
01 Retrieval: Relevance Does Not Make Good Evidence
Complex questions need a complete, complementary evidence set with little duplication. Ranking documents by individual relevance rarely produces one. AdaTutoRank reframes reranking as evidence-set construction. Its three-level, nine-dimension rubric supplies both cold-start silver labels and reinforcement-learning rewards.
Conventional set scores let redundant documents share the reward and make essential evidence share the penalty. AdaTutoRank instead adjusts tutoring intensity to candidate-set quality. Strong outputs are evaluated only against the rubric. Weaker outputs receive a reference set and, when needed, a reflection on the differences.
A teacher model then rescores outputs with and without guidance. AdaTutoRank distills the improvement into per-token advantages, separating real evidence from free riders. The abstract reports the best overall results across ten RAG, deep-research, and set-evaluation benchmarks, with fewer retrieval calls. Exact gains and retrieval costs need confirmation from the full paper. The self-selector’s sibling sets and the self-reflector’s reflections come from the policy’s own frozen snapshot. This closed loop could amplify existing biases. Even so, binary relevance labels are no longer enough for training rerankers.
Key takeaways:
- Evaluate completeness, complementarity, and redundancy, not only individual relevance.
- Set-level scores hide essential evidence and free-rider documents, demanding finer credit assignment.
- Check whether snapshot-derived reference sets and reflections amplify bias inside the training loop.
02 Efficiency: Score Globally, Then Cut Compute
MALA does not prune attention connections in advance. It scores every valid causal interaction, then uses normalized attention contributions to decide whether later computation is worthwhile. This preserves global reach while skipping much of the low-value work after scoring. It offers a third route beyond dense attention and predefined sparsity.
In 128K-context operator tests, the abstract reports 2.2× faster training forward passes and 3.0× faster backward passes. Decoding improved by 1.6×. At 8K, associative-recall accuracy reached 89.67%, close to full attention’s 89.97%.
Training runs from 0.6B to 14B largely matched full attention on perplexity while reducing total training FLOPs. Contribution-based compute allocation therefore appears able to scale. Operator latency, total training compute, and end-to-end application latency are different metrics. The last still needs full-system validation.
Key takeaways:
- Preserve every valid attention score, then reduce later computation according to its actual contribution.
- Long-context operator speedups are large while task accuracy stays close to full attention.
- Measure end-to-end latency instead of applying operator-level multipliers directly.
Source: MassAlloc Attention: Let Attention Allocate Its Own Compute
03 Multimodal: Make the Teacher Follow the Student
Choosing the best 5% of visual tokens is not the hardest part. With incomplete input, the model follows a different generation trajectory. Standard distillation may never teach it how to recover from those states.
LT-OPD first lets the low-token student generate. A frozen full-token copy then provides distribution-level supervision along the student’s actual trajectory. The teacher corrects the student where it wanders instead of only demonstrating an ideal route. A curriculum gradually reduces the visual-token budget to keep training stable.
Across nine Qwen3.5-4B benchmarks, average retained performance rose from 68.6% to 82.3% with only 5% of visual tokens. The method also transferred to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. KV-cache use and prefill compute both fell by about 85%, with no added inference overhead. Extreme compression must address the behavioral distribution shift after pruning, not only token selection.
Key takeaways:
- At a 5% visual-token budget, the student’s own generation trajectory becomes essential training data.
- A full-token teacher should correct the student along deployed states, not only teach ideal paths.
- Evaluate both token selection and post-compression distribution matching.
Source: Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction

Also Worth Noting
Today's Observation
The three featured papers share a finer-grained approach to efficiency, but they act at different stages. AdaTutoRank allocates teaching intensity by candidate-set quality during training. LT-OPD restores supervision along the low-token model’s own generation states. MALA allocates later operator compute after scoring every valid attention pair.
Extra training cost, operator speedups, and downstream quality retention need separate accounting. Percentages from unrelated benchmarks are not comparable.
Start by locating where compression losses concentrate: candidate sets, generation states, or attention connections. In the next compression experiment, log the quality loss for each reduced unit and its stage. Target the highest-loss group first.