AI Research Brief
Search
Methodology
中文
User Code Edits Cut Agent Success by 7.7 Points
9 selected from 200 papers
Featured
UEmbed: Unified Sparse and Dense Multimodal Embeddings
score 10
入选 HF Daily Papers; HF 热度: 43 upvotes (+4); 有代码实现; 关键词(2): retrieval-augmented, agentic
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
score 10
入选 HF Daily Papers; HF 热度: 23 upvotes (+4); 有代码实现; 关键词(1): coding
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning
score 10
入选 HF Daily Papers; HF 热度: 22 upvotes (+4); 有代码实现; 关键词(2): scaling, reasoning
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
score 9
入选 HF Daily Papers; HF 热度: 30 upvotes (+4); 有代码实现
GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation
score 9
入选 HF Daily Papers; 有代码实现; 关键词(1): finetuning; 顶会接收: ECCV
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
score 8
入选 HF Daily Papers; HF 热度: 28 upvotes (+4); 关键词(3): scaling, fine-tuning, synthetic data
ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step
score 9
入选 HF Daily Papers; HF 热度: 10 upvotes (+3); 有代码实现; 关键词(1): reasoning
Also Worth Noting
Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration
score 4
关键词(2): deployment, robotics; 顶会接收: ECCV
Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees
score 3
顶会接收: ICLR