reinforcement-learning
28 talks
Einstein Arena: Harnessing Collective Agent Intelligence for Open Science
Trading Desks to Clinical Trials: Parallels in Applied Vertical AI
Training Krea 2: What Matters in Generative Model Training
From RL to IRL
Scaling Compute on Context
Scaling up Continual Learning
Taking Reinforcement Learning Cross Datacenter
Local Models: Trust, Control, Optimization
Teaching AI to Find Real Vulnerabilities
Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It
Data and Environment Curation for Post-Training LLMs
Ending AI Slop
Learning on the Job: The Future of Post-Training
Reinforcement Learning without Verifiable Rewards
Scaling to Long Horizons
The Base Model Is Dead
What's Next After RLHF?
Why Off-the-Shelf AI Doesn't Understand Money
Everything Is a Rollout
Harness Engineering is Not Enough: Why Software Factories Fail
Local Agentic Theory For Mobile Games
Special Topics in Kernels, RL, Reward Hacking in Agents
Computer-Use 2.0: Agents Just Got Multi-Cursor
Recursive Model Improvement
Modern Post-Training: A Deep Dive
How We Taught Agents to Use Good Retrieval
Continual Learning for AI Agents: From Failures to Durable Improvements
Stop Making Models Bigger, Make Them Behave