Today's Overview
- Agentic Retrieval Trades 160× Latency for 8.7 nDCG@10 Points. With the same embedding model, average query latency rose from 0.67 to 107.4 seconds while consuming 764,100 input tokens. It may be best reserved for high-value, low-frequency work.
- OnePO Turns the Teacher Into Temporary Guidance. It first reinforces unlikely but informative teacher tokens, then removes guidance once the current policy beats the teacher. Training on 20K medical samples produced a 67.2 HealthBench Total score. That beat SFT+RL by 2.7 points and pure RL by 7.4 points. The reported evidence is limited to medical adaptation.
- ARRO Penalizes Both Missed and Unwanted Edits. A vision-language reward model audits candidate edits and generates rewards. On FLUX.1 Kontext-dev, average EditScore rose from 5.21 to 5.88, while unwanted pixel changes fell 8.4% across 600 samples.
Featured
01 Retrieval: Are 8.7 Points Worth 107 Seconds?
On average, agentic retrieval took 107.4 seconds per query, versus 0.67 seconds for standard retrieval. It also consumed 764,100 input tokens per query. That is a steep bill for an 8.7-point nDCG@10 gain over standard retrieval using the same embedding model.
The method combines LLM reasoning with retriever-based corpus exploration inside a ReAct loop. With the same embedding model, the workflow produced competitive results on both ViDoRe v3 and BRIGHT. Its generalization may be stronger than specialized retrieval methods optimized for a single domain.
The roughly 160× average latency and 5.8K output tokens per query make direct replacement of high-volume search a poor fit under typical constraints. Reserve it for due diligence, complex research, and expensive decisions. Keep a fast retrieval path for ordinary queries, then validate costs against model pricing and expected concurrency.
Key takeaways:
- With the same embedding model, agentic retrieval can trade inference for an 8.7-point nDCG@10 gain over standard retrieval.
- The reported averages of 107.4 seconds and more than 760K input tokens make frequent-query deployment costly and operationally difficult.
- Route queries by value instead of replacing standard retrieval across the board.
Source: Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks
02 Training: OnePO Simplifies RL-Only Domain Adaptation
Teacher models often guide the entire adaptation process. Once the student becomes stronger, stale teacher outputs can anchor it to an outdated distribution.
OnePO treats the teacher as a temporary coach. Early training reinforces low-probability teacher tokens that carry useful information. This addresses the shortage of effective gradients during RL-only cold starts.
Guidance ends when the current policy surpasses the teacher output, avoiding Teacher-Distribution Anchoring. With 20K medical samples, OnePO scored 67.2 on HealthBench Total. It beat SFT followed by RL by 2.7 points and pure RL by 7.4 points. Teams can test this simpler path, but should not assume the medical gains transfer to other domains.
Key takeaways:
- In OnePO's medical-adaptation experiments, transient teacher guidance addressed weak early gradients and later Teacher-Distribution Anchoring.
- Domain RL must address both weak early gradients and later anchoring to the teacher's distribution.
- OnePO could simplify SFT-plus-RL pipelines, but the reported evidence is limited to medical adaptation.
Source: HuatuoGPT-3: RL-Only Domain Adaptation from Base Models
03 Image Gen: Edit Only What Users Ask
Image editors can miss requested changes or alter unrelated content. ARRO trains against both failures instead of measuring instruction compliance alone.
A vision-language reward model audits candidate edits for omissions and unwanted changes, then generates the training reward. The process aggregates and verifies problems across multiple candidates. It turns minimal-change editing into a scalable, two-sided constraint without per-instruction human labels.
On FLUX.1 Kontext-dev, average EditScore climbed from 5.21 to 5.88. Unwanted pixel changes fell 8.4% across 600 samples. The method also improved OmniGen2, letting product teams train for both compliance and restraint.
Key takeaways:
- Editing evaluations must measure both missed changes and unwanted changes.
- Vision-language models can create scalable rewards when per-instruction human labels are unavailable.
- Minimal-change behavior can become a directly optimized product constraint.
Source: Scalable Minimal-Change Learning for Controllable Image Editing

Also Worth Noting
Today's Observation
All three featured papers replace fixed pipelines with feedback loops that respond to current conditions. The retrieval agent chooses its next search. OnePO decides whether to retain guidance based on whether the current policy has surpassed the teacher. ARRO generates rewards from omissions and unwanted edits in candidate results.
These studies do not prove that adaptive loops generally outperform fixed workflows. They do set a stricter acceptance bar. Track iteration counts, exit conditions, and resource costs such as tokens, latency, and GPU memory. Measure missed and unwanted edits separately so gains cannot hide loop delays, stale guidance, or one-sided metrics.
Start by giving these four data groups separate fields in evaluation dashboards. Then compare their distributions across successful and failed tasks.