TRM Defines Rubrics; SAKI Reaches 4.22× Throughput

Today's Overview

  • TRM Defines Rubrics Before Assigning Fine-Grained Scores. PD-GRPO reduces the score polarization caused by conventional pairwise preference optimization, giving visual generation RL more specific reward signals.
  • SAKI Switches Supervision Based on Teacher-Student Disagreement. It preserves reverse KL for accepted positions and directly learns the teacher’s preferred token at correction points. Generation throughput reaches 4.22× on matching workloads.
  • GEB Tracks Long-Video Memory Through Entity Identity. It improves results across four day-long to week-long video benchmarks. EgoLifeQA accuracy reaches 72.0%, beating the previous best by 4.4 percentage points.

Featured

01 Score Images Against Case-Specific Rubrics

A single score cannot explain whether an image failed on subject consistency, text accuracy, or edit boundaries. TRM changes the order. It first creates a rubric for each generation or editing case, then checks individual criteria and assigns fine-grained scores.

PD-GRPO addresses a separate problem. Conventional pairwise preference optimization can push scores toward extremes. The method keeps pointwise scoring while using pairwise supervision to improve discrimination.

The abstract reports that TRM beats other open reward models on image generation and editing benchmarks. It also competes with closed alternatives. Using TRM as the reward in reinforcement learning consistently improves diverse visual generation models, indicating that its fine-grained, case-adaptive rewards provide effective optimization signals. Tests on harder real-world prompts must still confirm whether the rubrics stay reliable.

Key takeaways:

  • Case-specific criteria make complex visual evaluations more informative than a direct aggregate score.
  • Keep the two contributions distinct: rubrics define what to judge, while PD-GRPO reduces score polarization from pairwise training.
  • Visual generation teams should test whether fine-grained rewards provide steadier optimization than a single scalar.

02 Route Supervision Through Disagreement

When the teacher accepts a student-sampled token, reverse KL preserves information from the online trajectory. Direct teacher-token supervision activates only when the two distributions genuinely disagree and correction becomes necessary.

SAKI makes this disagreement explicit through maximal coupling. The correction probability exactly equals the total variation distance between teacher and student distributions. A single trust-region constraint therefore controls trajectory drift, intervention frequency, and special supervision frequency.

Its speculative verifier preserves the target trajectory distribution while raising generation throughput to 4.22× on matching workloads. That is engineering support, not the method’s central contribution. Across seven mathematical reasoning benchmarks, both 0.6B and 1.7B students beat their corresponding teacher-guided baselines. Transfer to coding and tool use still needs testing.

Key takeaways:

  • Online distillation can route supervision according to teacher-student disagreement instead of treating every token identically.
  • Maximal coupling turns intervention frequency into a measurable and controllable training signal.
  • Evaluate supervision-driven quality gains separately from throughput gains delivered by speculative verification.

03 Video Memory Must Track Identity

Long-video question answering may fail even after retrieving the right event. The system must also know whether similar objects across different clips are the same physical entity.

Grounded Entity Biographies connect observations of one entity across clips into a searchable biography. Each observation retains its local context. During question answering, the model reads both event evidence and entity biographies, following identity clues instead of searching only by description.

GEB improves results across four benchmarks containing day-long to week-long recordings. EgoLifeQA accuracy reaches 72.0%, beating the previous best by 4.4 percentage points. Ablations show that adding more descriptions cannot fully replace entity linking. Evaluation should still test identity errors under occlusion, appearance changes, and scenes crowded with similar objects.

Key takeaways:

  • Retrieving the right event does not guarantee that a long-video system found the right object.
  • Entity-level memory suits applications that track items, people, or equipment across long periods.
  • Evaluate identity-linking errors separately instead of relying only on final question-answering accuracy.
TRM Defines Rubrics; SAKI Reaches 4.22× Throughput

Also Worth Noting

04
Complementary Trajectories Can Repair All-Failure Training Groups. TrainingGRAFT replaces all-failure groups, then controls cross-model bias through sequence-compatibility weighting and token-level importance-ratio clipping. link.
05
Fuse Multi-Level Visual Detail Without Expanding Latents. Image GenHiRAE anchors on deep representations and adds residual information from each layer under a norm budget, improving reconstruction and text-to-image alignment. link.
06
Validate Actions Before Teaching the Robot. RoboticsReal2Gym verifies demonstrations or retargeted actions through native physics execution. It converts successes and failures into procedures, relative object motions, and recovery strategies without updating base-model weights. link.
07
Explore Where Models Might Change Their Minds. TrainingHDL identifies branch points through token-likelihood changes caused by hindsight feedback. It samples alternative continuations only from those positions, focusing exploration while generating fewer tokens. link.
08
Agents Can Act Before Finishing Their Reasoning. AgentActFirst-OPD separates environment actions from full reasoning generation. It advances the interaction first, then produces thought-action responses for teacher supervision asynchronously. link.
09
Distill Each Video Capability Once Per Backbone Family. Video GenLongLive-Plug distills single-pass CFG, few-step sampling, and long-context correction into reusable LoRAs for compatible downstream models. link.
10
Geometry Code Can Scale Spatial Query Templates. MultimodalExemplar2VQA lets collaborating agents call geometry tools, turning varied static and object-centric query templates into large synthetic 3D question-answering datasets. link.