AI Research Brief
Search
Methodology
中文
Agents Score 58.1%, but Only 2.8% Pass Every Test
13 selected from 273 papers
Featured
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
score 10
入选 HF Daily Papers; HF 热度: 25 upvotes (+4); 有代码实现; 关键词(2): coding, open-source
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
score 8
入选 HF Daily Papers; HF 热度: 21 upvotes (+4); 关键词(5): scaling, GRPO, agentic, coding, open-source
GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills
score 6
入选 HF Daily Papers; HF 热度: 3 upvotes (+1); 有代码实现
Also Worth Noting
MintAct: A Unified Visual Agent for Digital Environments
score 4
入选 HF Daily Papers; 关键词(3): serving, tool use, vision-language
When Does Reasoning Help in Machine Translation? A Hierarchical Analysis of LRM Reasoning Traces
score 4
关键词(1): reasoning; 顶会接收: EMNLP
PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR
score 4
关键词(1): open-source; 顶会接收: EMNLP
CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation
score 4
机构: Alibaba; 关键词(1): text-to-video
GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation
score 4
关键词(2): distillation, reasoning; 顶会接收: EMNLP
PointLAM: Local Attentive Mamba for Efficient Point-based 3D Object Detection
score 4
关键词(3): quantization, latency, mamba; 顶会接收: ECCV
Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts
score 4
关键词(1): lightweight; 顶会接收: EMNLP
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
score 3
入选 HF Daily Papers
Cube-Splat: High-Fidelity 360° Gaussian Splatting SLAM via Cubemap Factorization and Adjoint-Consistent Optimization
score 3
顶会接收: ECCV
From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers
score 3
顶会接收: EMNLP