AI Research Brief
Search
Methodology
中文
AgentWorld's Best Model Succeeds Only 52% of the Time
20 selected from 200 papers
Featured
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
score 9
入选 HF Daily Papers; HF 热度: 115 upvotes (+4); 有代码实现
Block Sparse Attention with Log-Linear Complexity
score 8
入选 HF Daily Papers; HF 热度: 20 upvotes (+4); 关键词(2): scaling, reasoning
InternW0-$Δ$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data
score 7
入选 HF Daily Papers; HF 热度: 17 upvotes (+3); 关键词(3): distillation, open-source, open source
Game Arena: Strategic LLM Evaluation in Competitive Environments
score 7
机构: DeepMind; 入选 HF Daily Papers; HF 热度: 4 upvotes (+1)
AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
score 6
入选 HF Daily Papers; HF 热度: 6 upvotes (+2); 关键词(1): open-source
Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning
score 5
入选 HF Daily Papers; HF 热度: 8 upvotes (+2)
Softmax Reparameterization for Output-Head Quantization
score 5
入选 HF Daily Papers; HF 热度: 2 upvotes (+1); 关键词(5): scaling, compression, quantization, latency, post-training
Also Worth Noting
Externalized CPDAG Summaries Improve LLM Causal Deduction
score 4
关键词(1): reasoning; 顶会接收: NeurIPS
Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods
score 4
关键词(1): scaling; 顶会接收: NeurIPS
Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?
score 4
关键词(4): scaling, lightweight, reasoning, vision-language; 顶会接收: NeurIPS
Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs
score 4
关键词(3): lightweight, fine-tuning, reasoning; 顶会接收: NeurIPS
Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding
score 4
机构: Baidu; 关键词(2): reasoning, open-source
Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents
score 4
机构: Peking University; 关键词(4): compression, distillation, serving, throughput
SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery
score 4
关键词(3): scaling, reasoning, vision-language; 顶会接收: NeurIPS
GraphWrit3R: End-to-End 3D Scene Graph Writing
score 4
机构: Stanford; 关键词(2): deployment, open-source
Mutable Transcripts: Mitigating Context Pollution through Editable Conversation State
score 3
顶会接收: NeurIPS
Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
score 3
顶会接收: NeurIPS
PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents
score 3
顶会接收: EMNLP
Beyond Empirical Support: Structured Outlier Generation via Sinkhorn Optimal Transport
score 3
顶会接收: NeurIPS
Online Learning via Learned Latent Bayesian Tracking
score 3
顶会接收: NeurIPS