Today's Overview
- Post-Training Can Improve One-Shot Accuracy While Shrinking Solution Coverage. Across 42 comparisons spanning four model families and three agent benchmarks, pass@1 and pass@K sometimes moved in opposite directions. With sufficient test-time budget, base models often surpassed their post-trained counterparts in pass@K solution coverage.
- Image Generators Can Simulate GUI Environments. AutoGUIWorld creates interaction traces without installing software. Fine-tuning raised the average OSWorld task score from 33.0% to 40.8%, but reliable state transitions, action targets, and quality filters remain essential.
- 4Director Turns Directorial Intent Into Executable 3D Constraints. It reuses a canonical mesh and applies rigid transformations frame by frame, improving cross-view consistency. The generative model still handles non-rigid changes.
Featured
01 Post-Training Can Shrink Solution Coverage
Post-training makes models more accurate on a single attempt. Give them more inference budget, however, and base models may solve more tasks. The study examined this trade-off across 42 cases involving 14 base/post-trained model pairs, four model families, and three agent benchmarks.
Post-training pushed tasks toward two extremes: reliably solved or never solved. It raised pass@1 and consistency at the cost of solution coverage. With sufficient test-time budget, base models often surpassed their post-trained counterparts on pass@K. The paper calls this loss in test-time scaling capacity the Sharpening Tax. Its authors say a few rollouts can estimate it, though the full paper must confirm that estimate's stability.
The authors also propose PTGS, which adjusts sampling temperature by task difficulty. During reinforcement learning in two agent environments, PTGS improved both one-shot success and repeated-sampling coverage over fixed temperatures. Agent teams should select training methods against both pass@1 and the pass@K that matches deployment budgets.
Key takeaways:
- Evaluate agent post-training with both pass@1 and deployment-relevant pass@K.
- Use the Sharpening Tax to detect lost test-time scaling capacity.
- Fixed sampling temperatures may be suboptimal. Task-adaptive sampling deserves attention.
Source: Sharpening Tax in Post-Training
02 AutoGUIWorld Uses Image Generators as GUI Simulators
Image generators can do more than draw interfaces. They can act as visual simulators for software environments. AutoGUIWorld first generates an initial interface. A planner then specifies atomic actions and expected changes, while the image generator updates each screenshot step by step.
This process creates training traces with action targets and state transitions without installing the underlying software. Fine-tuning Qwen3.5-35B-A3B on this data raised its average OSWorld task score from 33.0% to 40.8%. ScienceBoard task success rose from 14.0% to 32.2%. Synthetic experience transferred to real tasks.
Visual realism alone will not make this approach work. Each state transition must be credible, action targets must be grounded, and transition-level quality filters must screen the generated data. Image generators could become scalable environment infrastructure for GUI agents, but their interaction traces should not be treated as real by default.
Key takeaways:
- Image generators can expand environment coverage without installing or running the corresponding software environments.
- Evaluate state transitions and action targets before visual quality.
- Real-task gains justify investment, but quality filtering must be a core engineering capability.
Source: AutoGUIWorld: Image Generators as Visual World Models for GUI Agent
03 4Director Gives Video Models a Control Contract
4Director reconstructs each object from an input image as a canonical mesh. Users then specify rigid transformations for every frame. This turns directorial intent into executable 3D geometric constraints.
The system reuses the same mesh instead of regenerating hidden object structure in each frame. That improves shape consistency across viewpoints and avoids the limits of 2D trajectories for depth and rotation. It renders the controlled scene as a depth video. A Motion Adapter then fills in appearance, lighting, and non-rigid motion.
Training uses RealCOD-Rigid, a dataset with 20,774 annotated videos. The paper reports better visual quality, camera control, and object control than previous methods. Its explicit contract mainly covers rigid motion, though. The Motion Adapter still synthesizes non-rigid dynamics rather than controlling them through the prescribed rigid transformations.
Key takeaways:
- Canonical meshes and frame-level rigid transformations turn vague prompts into executable 3D constraints.
- Reusing one object mesh improves cross-view consistency without repeatedly guessing hidden structure.
- Test control limits under deformation, occlusion, and complex lighting before adopting the system.
Source: 4Director: Controlling Video World Models with Rigid 3D Geometry

Also Worth Noting
Today's Observation
The three featured papers expose two product axes that teams often conflate: coverage and constraint following. Coverage determines how many environment states or valid solutions a system can reach. Constraint following measures whether each output faithfully executes the user's intent.
AutoGUIWorld expands training-state coverage through synthetic environments. 4Director restricts output freedom with explicit geometry. Sharpening Tax warns that post-training can improve one-shot reliability while reducing solution coverage.
Report state or solution coverage separately from constraint-following rates. Mark where each training, sampling, or generation stage expands possibilities and where it narrows them. The next evaluation cycle should track both metric groups independently.