quantization
27 talks
The Frontier AI Inference Cloud for Agents
Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads
What's New in Inference Engineering
Deep dive on LLM Inference at Scale
Compression at the Edge
Why Large? Tiny LMs and Agents on Edge and Robotics
Special Topics in Kernels, RL, Reward Hacking in Agents
Frontier Results, On Device
You Might Not Need 50 Diffusion Steps
Run Frontier AI at Home
Accelerating AI on Edge
TLMs: Tiny LLMs and Agents on Edge Devices with LiteRT-LM
Running LLMs on your iPhone: 40 tok/s Gemma 4 with MLX
Running LLMs Locally: Practical LLM Performance on DGX Spark
Serving Voice AI at $1/hr: Open-Source, LoRAs, Latency, Load Balancing
Reinforcement Learning, Kernels, Reasoning, Quantization & Agents
360Brew: LLM-based Personalized Ranking and Recommendation
Teaching Gemini to Speak YouTube: Adapting LLMs for Video Recommendations to 2B+DAU
Optimizing Inference for Voice Models in Production
RAG at scale: production-ready GenAI apps with Azure AI Search
Mastering LLM Inference Optimization From Theory to Cost Effective Deployment
A Practical Guide to Efficient AI
Everything You Need to Know About Fine-tuning and Merging LLMs
From Model Weights to API Endpoint with TensorRT-LLM
Low Level Technicals of LLMs
Harnessing the Power of LLMs Locally
AI Engineering 201: Inference