Qwen3.5-4B’s Nine-Benchmark Average Retention Reaches 82.3% at 5% Visual Tokens

Today's Overview

  • Relevant Documents Are Not a Good Evidence Set. AdaTutoRank adjusts tutoring intensity to candidate-set quality and distills guidance gains into per-token advantages, separating key evidence from “free-rider” documents.
  • MALA Scores Every Valid Attention Pair Before Allocating Compute. In 128K-context operator tests, it accelerated training forward passes by 2.2×, backward passes by 3.0×, and decoding by 1.6×.
  • Extreme Visual Compression Must Correct the Student’s Own Trajectory. With only 5% of visual tokens, LT-OPD raised Qwen3.5-4B’s average retained performance across nine benchmarks from 68.6% to 82.3%.

Featured

01 Retrieval: Relevance Does Not Make Good Evidence

Complex questions need a complete, complementary evidence set with little duplication. Ranking documents by individual relevance rarely produces one. AdaTutoRank reframes reranking as evidence-set construction. Its three-level, nine-dimension rubric supplies both cold-start silver labels and reinforcement-learning rewards.

Conventional set scores let redundant documents share the reward and make essential evidence share the penalty. AdaTutoRank instead adjusts tutoring intensity to candidate-set quality. Strong outputs are evaluated only against the rubric. Weaker outputs receive a reference set and, when needed, a reflection on the differences.

A teacher model then rescores outputs with and without guidance. AdaTutoRank distills the improvement into per-token advantages, separating real evidence from free riders. The abstract reports the best overall results across ten RAG, deep-research, and set-evaluation benchmarks, with fewer retrieval calls. Exact gains and retrieval costs need confirmation from the full paper. The self-selector’s sibling sets and the self-reflector’s reflections come from the policy’s own frozen snapshot. This closed loop could amplify existing biases. Even so, binary relevance labels are no longer enough for training rerankers.

Key takeaways:

  • Evaluate completeness, complementarity, and redundancy, not only individual relevance.
  • Set-level scores hide essential evidence and free-rider documents, demanding finer credit assignment.
  • Check whether snapshot-derived reference sets and reflections amplify bias inside the training loop.

02 Efficiency: Score Globally, Then Cut Compute

MALA does not prune attention connections in advance. It scores every valid causal interaction, then uses normalized attention contributions to decide whether later computation is worthwhile. This preserves global reach while skipping much of the low-value work after scoring. It offers a third route beyond dense attention and predefined sparsity.

In 128K-context operator tests, the abstract reports 2.2× faster training forward passes and 3.0× faster backward passes. Decoding improved by 1.6×. At 8K, associative-recall accuracy reached 89.67%, close to full attention’s 89.97%.

Training runs from 0.6B to 14B largely matched full attention on perplexity while reducing total training FLOPs. Contribution-based compute allocation therefore appears able to scale. Operator latency, total training compute, and end-to-end application latency are different metrics. The last still needs full-system validation.

Key takeaways:

  • Preserve every valid attention score, then reduce later computation according to its actual contribution.
  • Long-context operator speedups are large while task accuracy stays close to full attention.
  • Measure end-to-end latency instead of applying operator-level multipliers directly.

03 Multimodal: Make the Teacher Follow the Student

Choosing the best 5% of visual tokens is not the hardest part. With incomplete input, the model follows a different generation trajectory. Standard distillation may never teach it how to recover from those states.

LT-OPD first lets the low-token student generate. A frozen full-token copy then provides distribution-level supervision along the student’s actual trajectory. The teacher corrects the student where it wanders instead of only demonstrating an ideal route. A curriculum gradually reduces the visual-token budget to keep training stable.

Across nine Qwen3.5-4B benchmarks, average retained performance rose from 68.6% to 82.3% with only 5% of visual tokens. The method also transferred to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. KV-cache use and prefill compute both fell by about 85%, with no added inference overhead. Extreme compression must address the behavioral distribution shift after pruning, not only token selection.

Key takeaways:

  • At a 5% visual-token budget, the student’s own generation trajectory becomes essential training data.
  • A full-token teacher should correct the student along deployed states, not only teach ideal paths.
  • Evaluate both token selection and post-compression distribution matching.
Qwen3.5-4B’s Nine-Benchmark Average Retention Reaches 82.3% at 5% Visual Tokens

Also Worth Noting

04
CIS Tightens Importance-Weight Caps by Token Confidence TrainingIt targets probability mismatches between RLVR training and inference engines. Across five math-reasoning benchmarks, it beat evaluated baselines on average for all three mixture-of-experts models. link
05
KV Heads Can Share Complete Causal Coverage ArchitectureCoWA assigns different parts of long-range history to different heads. In 128K operator tests, training forward and backward latency fell to 1/7.4 and 1/8.6 of full attention. link
06
Meaning-Preserving Transformations Expose Asymmetric Generalization SafetyModels largely retain task performance after adapting to a transformation, but harmful-output rates can rise sharply. Alignment appears more dependent on the input’s surface distribution. link
07
Solving Every Intermediate Step Does Not Ensure the Answer ReasoningAcross 354 NuminaMath problems and six models, OracleLadder found composition gaps were the largest failure category for every model. link
08
External Memory Does Not Consistently Beat Native Memory AgentSCLATE compresses cross-month scenarios into hours with a shared event clock. Results vary widely across models using identical harnesses and memory configurations. Post-training Qwen3.5-4B improved without modifying either component. link
09
NTRL-Code Self-Improves From Noisy Unlabeled Instructions Code IntelligenceIt creates semantic anchors through conservative denoising, then builds proxy objectives by aggregating candidate programs through their abstract syntax trees. link
10
Models Improve Weak Starters More Consistently Than Editable Competent Baselines EvaluationOptiArena fixes five editing rounds, scaffolding, and resource budgets. It also checks whether editing degrades competent baselines. link
11
Delayed QA Supervision Can Train Test-Time Memory ArchitecturePlacing questions after many unrelated events directly trains long-term fact retention and updates. How much of the gain comes specifically from the delay still needs isolation. link
12
World SLAM Model Turns SLAM Into Long-Horizon Memory RoboticsIt brings persistent state updates, durable memory, and back-end error correction into end-to-end navigation instead of treating SLAM as an upstream module. link
13
Ad-Account Detection Agents Should Investigate, Not Decide AgentIn compromised business ad-account detection, a neuro-symbolic stage makes the final decision using symbolic rules, naive Bayes calibration, and a data-tuned contradiction layer. Precision rose from 0.250 to 0.446, while recall fell from 0.920 to 0.660. link
14
PULSE Selects Demonstrations by Internal Utility Features InterpretabilityIt finds sparse-autoencoder features associated with gains for the target model, replacing retrieval based only on textual similarity. link
15
Ignored Curvature Couplings Cause Low-Rank Collapse TrainingAt high compression rates, SDLRT restores important singular directions with a lightweight compensation buffer. Negative feedback then stabilizes each layer’s rank. link

Today's Observation

The three featured papers share a finer-grained approach to efficiency, but they act at different stages. AdaTutoRank allocates teaching intensity by candidate-set quality during training. LT-OPD restores supervision along the low-token model’s own generation states. MALA allocates later operator compute after scoring every valid attention pair.

Extra training cost, operator speedups, and downstream quality retention need separate accounting. Percentages from unrelated benchmarks are not comparable.

Start by locating where compression losses concentrate: candidate sets, generation states, or attention connections. In the next compression experiment, log the quality loss for each reduced unit and its stage. Target the highest-loss group first.