inference
19 talks
KV Cache-Aware Routing and P/D Disaggregation on Kubernetes
Preferences Over Benchmarks: Model Routing
Generative Video at the Speed of Light
The Next Medium: Why Real-Time Interactive Video Changes Everything
Voice agents with Realtime Video
Taking Reinforcement Learning Cross Datacenter
Compression at the Edge
The State of Model Routing
Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It
Why Large? Tiny LMs and Agents on Edge and Robotics
Stop Renting Your Cognitive Infrastructure
Special Topics in Kernels, RL, Reward Hacking in Agents
State of the Union: Why Local, Why Now
OpenClaw in Your Hand: Building a Physical AI Terminal
You Might Not Need 50 Diffusion Steps
Sovereign Escape Velocity: Ownership with Open Models
GPU Cloud Deployment Without Leaving Your IDE
Road to 5 Million Tokens: Breaking Barriers in Long Context Training
Under 5 minutes to a deployed LLM endpoint