Agent 124 briefings On-Policy Feedback Makes Training Signals Useful Voice RAG’s F1 Drop Widens by 67% Deployment Feedback Speeds Up Inference by 6.98× View topic →
Multimodal 123 briefings 300 Tasks Test Visual Reasoning, Retrieval Runs 12.4× Faster On-Policy Feedback Makes Training Signals Useful Voice RAG’s F1 Drop Widens by 67% View topic →
Evaluation 120 briefings 300 Tasks Test Visual Reasoning, Retrieval Runs 12.4× Faster On-Policy Feedback Makes Training Signals Useful Voice RAG’s F1 Drop Widens by 67% View topic →
Training 104 briefings On-Policy Feedback Makes Training Signals Useful Voice RAG’s F1 Drop Widens by 67% Executable Rubrics Cut Evaluation Latency by Up to 320× View topic →
Image Gen 98 briefings Training Bigger Models Without Paying the Full Price Life Agents, 4D Video, and Sparse Domain Updates Runtime Control Moves Into the Open View topic →
Safety 93 briefings 300 Tasks Test Visual Reasoning, Retrieval Runs 12.4× Faster On-Policy Feedback Makes Training Signals Useful Executable Rubrics Cut Evaluation Latency by Up to 320× View topic →
AI for Science 89 briefings 300 Tasks Test Visual Reasoning, Retrieval Runs 12.4× Faster On-Policy Feedback Makes Training Signals Useful Executable Rubrics Cut Evaluation Latency by Up to 320× View topic →
Efficiency 80 briefings On-Policy Feedback Makes Training Signals Useful Voice RAG’s F1 Drop Widens by 67% Deployment Feedback Speeds Up Inference by 6.98× View topic →
Architecture 76 briefings Executable Rubrics Cut Evaluation Latency by Up to 320× Deployment Feedback Speeds Up Inference by 6.98× 3D World-Building Success Stays Below 60% View topic →
Video Gen 74 briefings Voice RAG’s F1 Drop Widens by 67% Thirteen Skills Push Hallucination Detection to 92% Life Agents, 4D Video, and Sparse Domain Updates View topic →
Robotics 73 briefings Editing Live Video, Shrinking VLMs to 2.7 Bits Training Bigger Models Without Paying the Full Price Learning Workflows Lift Design Agents Across Seven Setups View topic →
Reasoning 63 briefings Voice RAG’s F1 Drop Widens by 67% Editing Live Video, Shrinking VLMs to 2.7 Bits Training Bigger Models Without Paying the Full Price View topic →
Retrieval 63 briefings 300 Tasks Test Visual Reasoning, Retrieval Runs 12.4× Faster On-Policy Feedback Makes Training Signals Useful Voice RAG’s F1 Drop Widens by 67% View topic →
Interpretability 60 briefings Voice RAG’s F1 Drop Widens by 67% Executable Rubrics Cut Evaluation Latency by Up to 320× Thirteen Skills Push Hallucination Detection to 92% View topic →
Code Intelligence 36 briefings On-Policy Feedback Makes Training Signals Useful A 4B Agent Ties the Giants by Verifying, Not Searching Harder The Model as Its Own Teacher, and Why Splitting Crowds Wins View topic →