
Why first: Frye supplies the vocabulary for the rest of the pack: tensor operations, memory bandwidth, batching, cold starts and the choice between an API and self-hosting. His broad map reaches the GPU, where the next talk looks more closely at why some inference workloads keep the hardware busy and others leave it waiting.








