Today's Overview
- Bigger Teachers Do Not Guarantee Better Distillation Feedback. Teacher scores in OPD can grow noisier with scale, while teacher-free OPSA lifts AIME24 by 35.41 points.
- PaperGym Turns Research-Plan Review Into an Auditable Reward Environment. Separating problem and rubric sources cuts rubric leakage to 3.7%.
- CAST Uses Action-Level Critiques to Reduce Costly Tool-Calling Errors. It beats GPT-OSS-120B by over 10% on consecutive retail-task passes and transfers to telemedicine.
- NoisEasier Fixes Video Object Relationships Without Fine-Tuning. Test-time diffusion-noise optimization improves difficult dimensions such as attribute binding and object interaction by over 10% on average.
Featured
01 Bigger Teachers Can Give Noisier Feedback
A larger teacher does not always provide more trustworthy feedback. In on-policy distillation (OPD), the teacher scores student outputs that are off-policy from its own perspective. The paper reports substantial noise in these scores, with the noise rate increasing alongside teacher size.
More surprisingly, keeping or removing the noisy signals produces similar student performance. Learning appears driven less by transferred teacher knowledge than by repeatedly suppressing low-probability tokens. A fixed negative advantage can even match the original method.
That result leads to OPSA, which needs no teacher supervision. It applies strong signals where model entropy is high, suppressing tail tokens and reallocating probability toward likely choices. Against base Qwen3-1.7B, it raises AIME24 Avg@32 by 35.41 points, a 263% relative gain. It also beats OPD by 16.77 points and more than doubles Pass@32 across three benchmarks. These are multi-sample metrics, so the full paper must clarify whether single-answer quality also improves.
Key takeaways:
- Do not equate a larger teacher with cleaner supervision. Measure its noise on student trajectories.
- Add fixed negative-advantage and teacher-free ablations before crediting gains to knowledge distillation.
- Check single-answer accuracy alongside Pass@32 to avoid overstating multi-sample gains.
Source: Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
02 PaperGym Makes Research Rewards Auditable
Research plans rarely have one correct answer, which makes reliable reinforcement-learning rewards difficult to construct. PaperGym converts papers into training tasks with inspectable evaluation criteria.
It generates questions from research goals and background sections. Rubrics come separately from methods and experiments. This separation makes paraphrasing the source less profitable and cuts rubric leakage to 3.7%. Existing datasets range from 11.90% to 34.10%.
Models first learn with rubric references, then receive GRPO rewards from the same criteria. Across Qwen3-1.7B, 4B, and 8B, average scores rise by 5.6, 5.0, and 4.8 points. The sequence also beats either stage alone and the reversed order. The remaining question is whether these rubrics capture genuine research quality across fields.
Key takeaways:
- Build a trustworthy review environment before expanding open-ended research-plan data.
- Extract questions and rubrics from different paper sections to reduce leakage and paraphrasing shortcuts.
- Audit whether rubric scores represent genuine research quality, not just benchmark improvement.
Source: PaperGym: Rubric-Centered Evolution for Research-Plan Generation
03 CAST Targets Costly Agent Mistakes
Long-horizon agents cannot be judged only by final task success. One irreversible action can make every later recovery attempt irrelevant.
CAST converts sparse task outcomes into action-level supervision. It synthesizes structured rationales that explain whether each action was valid. A critique model then turns those assessments into policy-optimization data.
After fine-tuning Qwen3 models, CAST beats GPT-OSS-120B by over 10% on four-run consecutive pass rates for retail tasks. It also delivers a 9% gain when transferred to unseen telemedicine scenarios. The approach offers a path to fewer costly local errors, though critique-model bias still needs closer examination.
Key takeaways:
- Use action-level supervision for high-risk tool calls instead of relying only on end-of-task rewards.
- Synthesize structured action rationales from agent trajectories and sparse outcomes.
- Track stability across steps and repeated runs, not just one-shot success rates.
Source: CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents
04 NoisEasier Optimizes Video Noise at Test Time
Fine-grained text alignment often depends on reward fine-tuning. New prompt distributions may require retraining, while optimization can invite reward hacking.
NoisEasier moves correction to inference time. It leaves model weights unchanged and optimizes diffusion noise with differentiable rewards. The method jointly adjusts the full stochastic generation trajectory, not just the initial noise. A short-step generator and multiple reward objectives keep costs under control.
The abstract reports consistent gains across several models. Difficult dimensions, including attribute binding, object interaction, and quantity understanding, improve by over 10% on average. Test-time optimization still adds inference cost, and further testing must verify that rewards track actual video quality.
Key takeaways:
- Consider weight-free test-time optimization when prompt distributions change frequently.
- Gains above 10% on difficult dimensions suggest that noise trajectories offer an effective control surface.
- Measure added inference costs and reward-hacking risks alongside alignment scores.
Source: NoisEasier: Test-Time Noise Optimization for Text-to-Video Generation

Also Worth Noting
Today's Observation
OPD, PaperGym, CAST, and PAC expose the same post-training bottleneck. Feedback must be trustworthy, independent, timely, and still useful for learning.
Dense teacher scores can be noisy on student trajectories. Rubrics derived from the same source as questions can leak answer clues. End-of-trajectory rewards arrive too late to prevent costly, irreversible actions such as incorrect refunds. Fixed task mixtures keep spending compute on mastered material.
These risks suggest four checks for every feedback pipeline. Can the evaluator judge data from the current policy reliably? Is the evaluation standard independent of the answer source? Does the signal identify the moment an error occurs? Do task weights change with learning progress?
Make these checks deployment gates before the next post-training run. Give each one measurable thresholds, fallback procedures, and a regular retesting schedule.