deployment
21 talks
KV Cache-Aware Routing and P/D Disaggregation on Kubernetes
The Agent Behind the Curtain: Building the Oz Cloud Agent Platform
How I automate my own job at Hugging Face using agents
Generative Video at the Speed of Light
Infra behind Krea 2: How to train and serve at scale
Taking Reinforcement Learning Cross Datacenter
Gadgets: Personal app vibe coding that is actually safe
What's Next After RLHF?
How Forward Deployed Engineering is Done at Cognition
How Forward Deployed Engineering is done at Kepler
Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub
Your LLM Stack Is a 2008 Database With Better Marketing
Agents Need Feature Flags
Stop Renting Your Cognitive Infrastructure
From fork() to Fleet: Designing an Agent Sandbox Cloud
State of the Union: Why Local, Why Now
The Pipeline Is Dead
Research to Reality: Bringing Frontier ML Research to Production
Sovereign Escape Velocity: Ownership with Open Models
GPU Cloud Deployment Without Leaving Your IDE
Under 5 minutes to a deployed LLM endpoint