AI Research Brief
Search
Methodology
中文
On-Policy Feedback Makes Training Signals Useful
19 selected from 320 papers
Featured
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
score 9
入选 HF Daily Papers; HF 热度: 22 upvotes (+4); 有代码实现
On-Policy Self-Distillation in Diffusion Models
score 10
入选 HF Daily Papers; HF 热度: 54 upvotes (+4); 有代码实现; 关键词(2): distillation, post-training
On-policy Distillation with Verifiable Reward
score 9
入选 HF Daily Papers; HF 热度: 15 upvotes (+3); 有代码实现; 关键词(4): distillation, GRPO, post-training, reasoning
TorchMorph: CUDA-accelerated Morphological Transforms
score 7
入选 HF Daily Papers; HF 热度: 3 upvotes (+1); 有代码实现; 关键词(2): lightweight, throughput
Also Worth Noting
CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
score 5
入选 HF Daily Papers; HF 热度: 4 upvotes (+1); 关键词(1): agentic
Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding
score 4
关键词(2): scaling, reasoning; 顶会接收: EMNLP
Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression
score 4
机构: University of Toronto; 关键词(6): compression, quantization, deployment, fine-tuning, post-training
When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs
score 4
关键词(1): synthetic data; 顶会接收: EMNLP
Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping
score 4
关键词(3): fine-tuning, code generation, coding; 顶会接收: EMNLP
MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation
score 9
入选 HF Daily Papers; 有代码实现; 关键词(1): deployment; 顶会接收: EMNLP
Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research
score 4
关键词(2): agentic, open-source; 顶会接收: EMNLP
Low-Rank Ternary Adaptation for Fine-Tuning Transformers
score 4
关键词(3): quantization, fine-tuning, fine-tune; 顶会接收: ECCV
IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves
score 4
关键词(1): distillation; 顶会接收: ECCV
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
score 4
关键词(2): GRPO, guardrails; 顶会接收: EMNLP
Meta$^n$: Recursive Self-Improvement through Emergent Depth
score 6
入选 HF Daily Papers; HF 热度: 10 upvotes (+3)
SENSESHIFT: Continuous Sentiment-Controlled Text Generation via Encoder-based Mask Infilling
score 3
顶会接收: EMNLP
Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning
score 3
顶会接收: ICML
Expectation, Backlash, Recovery, and Excitement: How Model Releases Shape Reddit Perceptions of Conversational AI Systems
score 3
顶会接收: EMNLP
Arbitrary Polygon Oscillator: Generalizing Polygonal Synthesis to Arbitrary Shapes, Morphing, and Three-Dimensional Polyhedra
score 3
机构: MIT