Today's Overview
- TRM Defines Rubrics Before Assigning Fine-Grained Scores. PD-GRPO reduces the score polarization caused by conventional pairwise preference optimization, giving visual generation RL more specific reward signals.
- SAKI Switches Supervision Based on Teacher-Student Disagreement. It preserves reverse KL for accepted positions and directly learns the teacher’s preferred token at correction points. Generation throughput reaches 4.22× on matching workloads.
- GEB Tracks Long-Video Memory Through Entity Identity. It improves results across four day-long to week-long video benchmarks. EgoLifeQA accuracy reaches 72.0%, beating the previous best by 4.4 percentage points.
Featured
01 Score Images Against Case-Specific Rubrics
A single score cannot explain whether an image failed on subject consistency, text accuracy, or edit boundaries. TRM changes the order. It first creates a rubric for each generation or editing case, then checks individual criteria and assigns fine-grained scores.
PD-GRPO addresses a separate problem. Conventional pairwise preference optimization can push scores toward extremes. The method keeps pointwise scoring while using pairwise supervision to improve discrimination.
The abstract reports that TRM beats other open reward models on image generation and editing benchmarks. It also competes with closed alternatives. Using TRM as the reward in reinforcement learning consistently improves diverse visual generation models, indicating that its fine-grained, case-adaptive rewards provide effective optimization signals. Tests on harder real-world prompts must still confirm whether the rubrics stay reliable.
Key takeaways:
- Case-specific criteria make complex visual evaluations more informative than a direct aggregate score.
- Keep the two contributions distinct: rubrics define what to judge, while PD-GRPO reduces score polarization from pairwise training.
- Visual generation teams should test whether fine-grained rewards provide steadier optimization than a single scalar.
Source: Think Before You Score: Thinking Reward Model for Visual Generation
02 Route Supervision Through Disagreement
When the teacher accepts a student-sampled token, reverse KL preserves information from the online trajectory. Direct teacher-token supervision activates only when the two distributions genuinely disagree and correction becomes necessary.
SAKI makes this disagreement explicit through maximal coupling. The correction probability exactly equals the total variation distance between teacher and student distributions. A single trust-region constraint therefore controls trajectory drift, intervention frequency, and special supervision frequency.
Its speculative verifier preserves the target trajectory distribution while raising generation throughput to 4.22× on matching workloads. That is engineering support, not the method’s central contribution. Across seven mathematical reasoning benchmarks, both 0.6B and 1.7B students beat their corresponding teacher-guided baselines. Transfer to coding and tool use still needs testing.
Key takeaways:
- Online distillation can route supervision according to teacher-student disagreement instead of treating every token identically.
- Maximal coupling turns intervention frequency into a measurable and controllable training signal.
- Evaluate supervision-driven quality gains separately from throughput gains delivered by speculative verification.
Source: SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
03 Video Memory Must Track Identity
Long-video question answering may fail even after retrieving the right event. The system must also know whether similar objects across different clips are the same physical entity.
Grounded Entity Biographies connect observations of one entity across clips into a searchable biography. Each observation retains its local context. During question answering, the model reads both event evidence and entity biographies, following identity clues instead of searching only by description.
GEB improves results across four benchmarks containing day-long to week-long recordings. EgoLifeQA accuracy reaches 72.0%, beating the previous best by 4.4 percentage points. Ablations show that adding more descriptions cannot fully replace entity linking. Evaluation should still test identity errors under occlusion, appearance changes, and scenes crowded with similar objects.
Key takeaways:
- Retrieving the right event does not guarantee that a long-video system found the right object.
- Entity-level memory suits applications that track items, people, or equipment across long periods.
- Evaluate identity-linking errors separately instead of relying only on final question-answering accuracy.
Source: Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
