AI论文简报
搜索
方法论
公众号
EN
用户改代码让Agent成功率降7.7个百分点
从200篇论文中选出9篇
重点关注
UEmbed: Unified Sparse and Dense Multimodal Embeddings
score 10
入选 HF Daily Papers;HF 热度: 43 upvotes (+4);有代码实现;关键词(2): retrieval-augmented, agentic
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
score 10
入选 HF Daily Papers;HF 热度: 23 upvotes (+4);有代码实现;关键词(1): coding
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning
score 10
入选 HF Daily Papers;HF 热度: 22 upvotes (+4);有代码实现;关键词(2): scaling, reasoning
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
score 9
入选 HF Daily Papers;HF 热度: 30 upvotes (+4);有代码实现
GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation
score 9
入选 HF Daily Papers;有代码实现;关键词(1): finetuning;顶会接收: ECCV
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
score 8
入选 HF Daily Papers;HF 热度: 28 upvotes (+4);关键词(3): scaling, fine-tuning, synthetic data
ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step
score 9
入选 HF Daily Papers;HF 热度: 10 upvotes (+3);有代码实现;关键词(1): reasoning
也值得关注
Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration
score 4
关键词(2): deployment, robotics;顶会接收: ECCV
Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees
score 3
顶会接收: ICLR