-
Language Models Can Control Their Own Attention
score 11
机构: MIT;入选 HF Daily Papers;HF 热度: 60 upvotes (+4);关键词(1): lightweight
-
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
score 10
入选 HF Daily Papers;HF 热度: 519 upvotes (+4);有代码实现;关键词(1): distillation
-
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
score 10
入选 HF Daily Papers;HF 热度: 113 upvotes (+4);有代码实现;关键词(3): lightweight, distillation, agentic
-
PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation
score 8
入选 HF Daily Papers;HF 热度: 6 upvotes (+2);有代码实现;关键词(2): code generation, coding
-
Cliff: Learning Process Rewards from the First Mistake
score 7
入选 HF Daily Papers;HF 热度: 16 upvotes (+3);关键词(4): distillation, GRPO, post-training, reasoning
-
CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
score 9
入选 HF Daily Papers;HF 热度: 8 upvotes (+2);关键词(1): real-time;顶会接收: EMNLP
-
Post-Training Language Models for Gold-Medal Performance in Coding Competitions
score 6
入选 HF Daily Papers;HF 热度: 9 upvotes (+2);关键词(4): fine-tuning, post-training, coding, reasoning
-
Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens
score 6
入选 HF Daily Papers;HF 热度: 2 upvotes (+1);有代码实现
-
Debias-SparseGPT: Bias-Aware Pruning for Large Language Models
score 8
入选 HF Daily Papers;HF 热度: 2 upvotes (+1);关键词(5): compression, quantization, pruning, deployment, post-training;顶会接收: EMNLP