AI Research Brief
Search
Methodology
中文
Eleven World Models Miss the Odds; ES Covers More Paths
19 selected from 334 papers
Featured
UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
score 10
入选 HF Daily Papers; HF 热度: 71 upvotes (+4); 有代码实现; 关键词(1): reasoning
What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
score 8
入选 HF Daily Papers; HF 热度: 59 upvotes (+4); 关键词(2): scaling, agentic
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
score 7
入选 HF Daily Papers; HF 热度: 15 upvotes (+3); 关键词(3): GRPO, post-training, reasoning
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
score 9
入选 HF Daily Papers; HF 热度: 82 upvotes (+4); 有代码实现
Also Worth Noting
Magpie: Real-Time World Renderer for Interactive Games
score 6
入选 HF Daily Papers; HF 热度: 7 upvotes (+2); 关键词(2): production, real-time
A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families
score 4
关键词(1): jailbreak; 顶会接收: EMNLP
Not Just Reason, Not Just Scan: Reinforcement Learning for Proactive Scientific Error Verification over Academic Paper
score 4
关键词(1): reasoning; 顶会接收: EMNLP
Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference
score 4
关键词(1): reasoning; 顶会接收: EMNLP
FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models
score 4
关键词(3): distillation, pruning, fine-tuning; 顶会接收: ICML
RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models
score 4
关键词(1): reasoning; 顶会接收: EMNLP
MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
score 4
关键词(3): lightweight, deployment, vision-language; 顶会接收: CVPR
From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
score 4
关键词(1): reasoning; 顶会接收: ECCV
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher
score 8
入选 HF Daily Papers; HF 热度: 67 upvotes (+4); 关键词(1): distillation
RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing
score 4
关键词(3): fine-tuning, GRPO, text-to-image; 顶会接收: EMNLP
Research Design Tracking and Assessment for the Social Sciences
score 4
关键词(1): RAG; 顶会接收: EMNLP
Prediction of Prediction (PoP): Inter-Layer Activation Fusion for Single-Pass Hallucination Detection in Large Language Models
score 4
机构: Mistral; 关键词(2): deployment, latency
Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition
score 4
关键词(1): vision-language; 顶会接收: EMNLP
BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing
score 4
关键词(3): scaling, agentic, reasoning; 顶会接收: EMNLP
MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE Framework
score 4
关键词(1): MoE; 顶会接收: ICML