latency
14 talks
Agentic Sites: Building Hyper Personalized Websites
Productionizing LLM Gateways: Architecture, Tradeoffs and Hard Lessons
KV Cache-Aware Routing and P/D Disaggregation on Kubernetes
Preferences Over Benchmarks: Model Routing
Generative Video at the Speed of Light
The Next Medium: Why Real-Time Interactive Video Changes Everything
How Web Data Infrastructure Powers the Next Generation of AI
Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub
Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing
Designing Voice Agents for Real Conversations
Your Voice Agent Doesn't Need a Frontier Model
The 100-Tool Agent Is a Trap
Voice In, Visuals Out: The Agony and the Ecstasy
Under 5 minutes to a deployed LLM endpoint