AgentWorld's Best Model Succeeds Only 52% of the Time

Today's Overview

  • Task Success Does Not Prove Good Multi-Agent Collaboration: AgentWorld tests 3–20 agents with asymmetric roles across 100 human-annotated tasks. The best model succeeds only 52.0% of the time. Its CCE metric traces causal dependencies between actions and measures how much team effort contributes to the outcome.
  • PISA Cuts Sparse Block Selection to (O(N\log N)): Hierarchical Top-K selection and fused Triton kernels avoid materializing the full score matrix. The abstract provides no end-to-end latency or memory results, so the engineering gains still need validation.
  • Natural-Image Priors Can Improve Elevation Maps: Researchers prune Stable Diffusion 3's text stream and condition it on photogrammetric DSMs and Pléiades imagery. In held-out Bordeaux, Dense Urban RMSE falls from 4.16 meters to 2.77 meters.

Featured

01 Success Does Not Equal Collaboration

A multi-agent system can finish a task without collaborating well. Binary success hides each member's contribution. It also cannot tell developers whether to fix communication, role assignment, or planning.

AgentWorld includes 100 human-annotated tasks and 100 scaled variants. Each requires 3–20 agents with asymmetric capabilities to collaborate for more than 50 consecutive rounds. Its Causal Collaboration Effectiveness metric traces dependencies between actions. CCE then measures how much team effort actually contributes to the result.

Even the best model reaches only a 52.0% task success rate. Common failures include broken communication, confused roles, and shared plans that decay across rounds. CCE offers a useful diagnostic approach, but the paper must establish whether its causal graphs are reliable and comparable across environments.

Key takeaways:

  • Do not evaluate agent teams using task success alone. Track whether member actions form a genuine chain of collaboration.
  • Test communication, role assignment, and shared planning as separate failure modes.
  • CCE offers a promising diagnostic method, but its causal judgments need further validation.

02 PISA Makes Block Selection Log-Linear

Sparse attention reduces attention computation, but block selection can hide a quadratic cost. Many systems still score every query against every candidate block before deciding what to keep.

PISA uses hierarchical Top-K selection to filter candidates from coarse to fine. Each level processes only a limited set. Pooling creates an (O(\log N))-level hierarchy, bringing total complexity to (O(N\log N)).

Fused Triton kernels handle hierarchical routing and LogSumExp scoring without materializing the full query-key score matrix. Language-modeling tests match baselines on commonsense reasoning and perform better on retrieval. The abstract omits end-to-end latency and memory usage, so real hardware tests must confirm the deployment gains.

Key takeaways:

  • Include block-selection cost when evaluating sparse attention. Counting only retained attention blocks gives an incomplete picture.
  • Hierarchical filtering provides an (O(N\log N)) implementation path for adaptive sparsity.
  • Do not infer deployment gains from theoretical complexity without end-to-end latency and memory measurements.

03 Elevation Refinement Needs Adapted Priors

Noise, outliers, and holes reduce the elevation accuracy of digital surface models from satellite stereo photogrammetry. Applying a natural-image diffusion model unchanged would ignore the structure of elevation data.

The researchers prune Stable Diffusion 3's text stream and add patch normalization. They condition the model on both photogrammetric DSMs and Pléiades satellite imagery. These changes adapt natural-image pretraining to elevation-map refinement.

In the sole held-out city, Bordeaux, Dense Urban RMSE drops from 4.16 meters to 2.77 meters. The adaptation method matters more than treating diffusion as an off-the-shelf geospatial tool. Results cover only the tested French cities and vertically aligned DSMs. They do not establish a general replacement for airborne LiDAR.

Key takeaways:

  • Natural-image diffusion priors can transfer to elevation refinement, but the model must reflect the data's structure.
  • Joint conditioning on DSMs and satellite imagery deserves more attention than either input alone.
  • Evaluate geographic generalization, alignment requirements, and actual cost boundaries against LiDAR.
AgentWorld's Best Model Succeeds Only 52% of the Time

Also Worth Noting

04
Chess, Poker, and Werewolf Test Strategy Across Extended Games EvaluationGame Arena covers perfect-information, imperfect-information, and multiplayer games. It tests planning, adaptation, and reliability under uncertainty. link
05
Equivalent Softmax Reparameterization Reduces W4 Quantization Error EfficiencyA separate result finds that quantizing Phi's output head cuts batch-one generation latency by 10.8% against a BF16 output-head baseline. link
06
Externalizing CPDAG State Improves Answers to Graph Queries ReasoningOn Corr2Cause's primary paired evaluation, typed and schema-constrained CPDAG summaries raise Qwen3.5-27B's (F_1)(Yes) from 73.0 to 86.4. link
07
Misaligned Audio-Video Links Drive Three-Modal Binding Failures MultimodalAdding existing detection boxes around active speakers improves results across four conversational-video benchmarks without training. link
08
Traceable Evidence and Query-Aware Summaries Cut Output Budgets RetrievalH2S-14B retains 97.1% of its 16K-budget performance with a 4K output budget. link
09
Compressing Only Older Tool Observations Creates a Tunable Tradeoff Code IntelligenceLOHA keeps the agent's actions and latest (K) observations verbatim. On SWE-bench Verified, larger (K) values generally trade less compression for higher solve rates. link
10
SatNav Builds 118,000 Drone Tasks Across 18 Cities RoboticsThree task types test loop-progress tracking, landmark-based localization, and route following with counting cues. The average trajectory spans 379 meters. link
11
Point Clouds and Gaussian Splats Become JSON Scene Graphs MultimodalGraphWrit3R avoids multi-stage pipelines that depend on ground-truth object annotations. It accepts point clouds, Gaussian Splats, or both. link
12
Hotel Price Sensitivity Varies More Than Tenfold Across Models AgentAcross 28 models, average booked room prices range from $247 to $393. Measure each purchasing agent directly; capability scores and model families do not predict its preferences. link
13
Low-Dimensional State Enables One-Step Bayesian Online Adaptation TrainingAURA learns a latent state space offline. It then uses an extended Kalman filter to update the state and reconstruct full parameters in non-stationary environments. link