AI论文简报
搜索
方法论
公众号
EN
Pivot-SD在LLaDA-8B-Instruct上以200题、每题4次采样胜过全序列SFT,Queen七轮从1782升至2697
从200篇论文中选出18篇
重点关注
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
score 12
机构: Stanford;入选 HF Daily Papers;HF 热度: 12 upvotes (+3);有代码实现;关键词(1): code generation
Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
score 11
机构: University of Toronto;入选 HF Daily Papers;HF 热度: 38 upvotes (+4);关键词(3): distillation, post-training, reasoning
Multilingual GSM-Symbolic: What determines capability transfer across languages?
score 10
入选 HF Daily Papers;HF 热度: 29 upvotes (+4);有代码实现;关键词(1): reasoning
World Embedding Benchmark
score 10
入选 HF Daily Papers;HF 热度: 28 upvotes (+4);有代码实现;关键词(2): lightweight, retrieval-augmented
ProAR: Learning Prospective Reasoning with Autoregressive Video Models
score 9
入选 HF Daily Papers;HF 热度: 18 upvotes (+3);有代码实现;关键词(3): lightweight, reasoning, embodied
Language Models that Play Chess and Explain Their Moves
score 9
入选 HF Daily Papers;HF 热度: 15 upvotes (+3);有代码实现;关键词(3): distillation, reasoning, robotics
Collective Bias Mitigation via Model Routing and Collaboration
score 9
机构: MIT;入选 HF Daily Papers;HF 热度: 10 upvotes (+3)
Native Action-Prior Learning from Videos for World Action Models
score 8
入选 HF Daily Papers;HF 热度: 62 upvotes (+4);关键词(1): pretraining
FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution
score 8
入选 HF Daily Papers;HF 热度: 15 upvotes (+3);有代码实现
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
score 7
入选 HF Daily Papers;HF 热度: 43 upvotes (+4)
也值得关注
Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT
score 4
机构: Cambridge;关键词(6): fine-tuning, DPO, PPO, GRPO, post-training
Geometry Meets Physics: Data-Efficient Pre-Training for Unstructured Neural PDE Solvers
score 4
关键词(2): fine-tuning, pre-training;顶会接收: NeurIPS
Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition
score 4
关键词(1): deployment;顶会接收: NeurIPS
Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis
score 4
机构: Cambridge;关键词(3): GRPO, coding, reasoning
Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain
score 4
关键词(1): scaling;顶会接收: NeurIPS
What Should World Models Forget? Stratified Retention for Continual Adaptation
score 4
关键词(2): latency, reasoning;顶会接收: NeurIPS
ZeroMAG: Zero-Shot Multimodal Adapter Generation for Plug-and-Play EEG Foundation Models
score 3
机构: Zhejiang University
Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis
score 3
顶会接收: NeurIPS