-
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
score 10
入选 HF Daily Papers; HF 热度: 25 upvotes (+4); 有代码实现; 关键词(2): reasoning, open-source
-
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
score 9
入选 HF Daily Papers; HF 热度: 32 upvotes (+4); 有代码实现
-
HelloWorld: Enabling Socially Interactive Characters in Video World Models
score 9
入选 HF Daily Papers; HF 热度: 18 upvotes (+3); 有代码实现; 关键词(1): distillation
-
FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory
score 9
入选 HF Daily Papers; HF 热度: 12 upvotes (+3); 有代码实现; 关键词(2): lightweight, compression
-
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
score 8
入选 HF Daily Papers; HF 热度: 63 upvotes (+4); 关键词(2): fine-tuning, GRPO
-
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
score 8
入选 HF Daily Papers; HF 热度: 52 upvotes (+4); 关键词(7): fine-tuning, GRPO, post-training, agentic, tool use
-
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
score 8
入选 HF Daily Papers; HF 热度: 52 upvotes (+4); 关键词(3): scaling, pretraining, MoE
-
The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
score 8
入选 HF Daily Papers; HF 热度: 38 upvotes (+4); 关键词(1): leaderboard
-
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap
score 9
入选 HF Daily Papers; HF 热度: 20 upvotes (+4); 有代码实现
-
K-EXAONE 2.0 Technical Report
score 7
入选 HF Daily Papers; HF 热度: 18 upvotes (+3); 关键词(6): post-training, pre-training, MoE, agentic, coding