Mean OSWorld Task Score Rises From 33.0% to 40.8%

Today's Overview

  • Post-Training Can Improve One-Shot Accuracy While Shrinking Solution Coverage. Across 42 comparisons spanning four model families and three agent benchmarks, pass@1 and pass@K sometimes moved in opposite directions. With sufficient test-time budget, base models often surpassed their post-trained counterparts in pass@K solution coverage.
  • Image Generators Can Simulate GUI Environments. AutoGUIWorld creates interaction traces without installing software. Fine-tuning raised the average OSWorld task score from 33.0% to 40.8%, but reliable state transitions, action targets, and quality filters remain essential.
  • 4Director Turns Directorial Intent Into Executable 3D Constraints. It reuses a canonical mesh and applies rigid transformations frame by frame, improving cross-view consistency. The generative model still handles non-rigid changes.

Featured

01 Post-Training Can Shrink Solution Coverage

Post-training makes models more accurate on a single attempt. Give them more inference budget, however, and base models may solve more tasks. The study examined this trade-off across 42 cases involving 14 base/post-trained model pairs, four model families, and three agent benchmarks.

Post-training pushed tasks toward two extremes: reliably solved or never solved. It raised pass@1 and consistency at the cost of solution coverage. With sufficient test-time budget, base models often surpassed their post-trained counterparts on pass@K. The paper calls this loss in test-time scaling capacity the Sharpening Tax. Its authors say a few rollouts can estimate it, though the full paper must confirm that estimate's stability.

The authors also propose PTGS, which adjusts sampling temperature by task difficulty. During reinforcement learning in two agent environments, PTGS improved both one-shot success and repeated-sampling coverage over fixed temperatures. Agent teams should select training methods against both pass@1 and the pass@K that matches deployment budgets.

Key takeaways:

  • Evaluate agent post-training with both pass@1 and deployment-relevant pass@K.
  • Use the Sharpening Tax to detect lost test-time scaling capacity.
  • Fixed sampling temperatures may be suboptimal. Task-adaptive sampling deserves attention.

02 AutoGUIWorld Uses Image Generators as GUI Simulators

Image generators can do more than draw interfaces. They can act as visual simulators for software environments. AutoGUIWorld first generates an initial interface. A planner then specifies atomic actions and expected changes, while the image generator updates each screenshot step by step.

This process creates training traces with action targets and state transitions without installing the underlying software. Fine-tuning Qwen3.5-35B-A3B on this data raised its average OSWorld task score from 33.0% to 40.8%. ScienceBoard task success rose from 14.0% to 32.2%. Synthetic experience transferred to real tasks.

Visual realism alone will not make this approach work. Each state transition must be credible, action targets must be grounded, and transition-level quality filters must screen the generated data. Image generators could become scalable environment infrastructure for GUI agents, but their interaction traces should not be treated as real by default.

Key takeaways:

  • Image generators can expand environment coverage without installing or running the corresponding software environments.
  • Evaluate state transitions and action targets before visual quality.
  • Real-task gains justify investment, but quality filtering must be a core engineering capability.

03 4Director Gives Video Models a Control Contract

4Director reconstructs each object from an input image as a canonical mesh. Users then specify rigid transformations for every frame. This turns directorial intent into executable 3D geometric constraints.

The system reuses the same mesh instead of regenerating hidden object structure in each frame. That improves shape consistency across viewpoints and avoids the limits of 2D trajectories for depth and rotation. It renders the controlled scene as a depth video. A Motion Adapter then fills in appearance, lighting, and non-rigid motion.

Training uses RealCOD-Rigid, a dataset with 20,774 annotated videos. The paper reports better visual quality, camera control, and object control than previous methods. Its explicit contract mainly covers rigid motion, though. The Motion Adapter still synthesizes non-rigid dynamics rather than controlling them through the prescribed rigid transformations.

Key takeaways:

  • Canonical meshes and frame-level rigid transformations turn vague prompts into executable 3D constraints.
  • Reusing one object mesh improves cross-view consistency without repeatedly guessing hidden structure.
  • Test control limits under deformation, occlusion, and complex lighting before adopting the system.
Mean OSWorld Task Score Rises From 33.0% to 40.8%

Also Worth Noting

04
Continuous Latents Become the Only Persistent State, While Discrete Tokens Add Scaffolding. ArchitectureHC-DLM targets both broken dependencies in parallel decoding and validity gaps in continuous representations. link
05
More Distinct Text Embeddings Are Not Always Better for Diffusion. TrainingSoft-label distillation brings plausible substitute words closer, creating a more connected latent space that is easier to sample. link
06
Long-Horizon Agents Need More Than Compressed History. AgentPoS explicitly tracks the current world, unresolved requirements, and stalled states. It checks belief consistency and selects recovery actions for different failure modes. link
07
Training Scenarios Should Change With Agent Failure Modes. AgentActiveSaddler's non-stationary curriculum improved test Pass@1 by 4.4 percentage points on GAIA2 and 7.5 points on Terminal-Bench 2.0. link
08
Fitting a Catalog Into Long Context Does Not Make It Usable. RetrievalRPTune first learns product ranking and pruning, then applies post-training with rewards relative to the current context. link
09
Execution Traces Can Test Whether Videos Follow Rules. EvaluationPROWBench uses program traces and timestamped events to move world-model evaluation from visual plausibility toward faithful execution. link
10
Long Audio-Video Reasoning Does Not Require Watching Everything. MultimodalOmniSeek decides when to look and when to listen. Its rewards discourage shortcuts that solve tasks through only one modality. link
11
Compiling Lean Code Does Not Prove the Right Claim. Code IntelligenceIsolated execution, statement comparison, assumption audits, and independent review turn successful outputs into traceable evidence. link
12
Hybrid Transformer-SSM Architectures Can Reuse μP Learning Rates. TrainingExperiments span widths from 256 to 2,048, depths from 4 to 32, and billion-parameter models. Empirical results currently run ahead of theory. link
13
VideoLLMs Learn Temporal Order, Then Lose It Across Layers. InterpretabilityRelevant representations peak in intermediate layers before fading. Reinjecting temporal-difference vectors at inference partially restores temporal reasoning without training. link
14
A Shared SE(3) Trajectory Representation Can Cover Objects, People, Cameras, and Robots. RoboticsWorld Motion Models use different masks to support prediction, completion, control, and cross-embodiment transfer. link
15
Open-Weight Models Still Struggle With Exact Kali Commands. EvaluationAcross 1,642 Kali tools, none exceeded 42% exact-command accuracy without explicit tool hints. KaliBench also provides verifiable training rewards without running commands. link

Today's Observation

The three featured papers expose two product axes that teams often conflate: coverage and constraint following. Coverage determines how many environment states or valid solutions a system can reach. Constraint following measures whether each output faithfully executes the user's intent.

AutoGUIWorld expands training-state coverage through synthetic environments. 4Director restricts output freedom with explicit geometry. Sharpening Tax warns that post-training can improve one-shot reliability while reducing solution coverage.

Report state or solution coverage separately from constraint-following rates. Mark where each training, sampling, or generation stage expands possibilities and where it narrows them. The next evaluation cycle should track both metric groups independently.