Exploring Next / Topics / Batching Topic Batching 1 episode Ep 879 Aug 18, 2026 Overview: Inference Optimization We finally give inference optimization its own episode — the idea that's quietly under half the stories we cover. We walk through what it actually means to make a trained model run faster and cheaper, from caching to quantization to batching, and why it matters more than almost anything else once a model ships. InferenceInference OptimizationAutoregressive GenerationKV Cache