Inference Engineering Is the Next Big Role in AI w/ Philip Kiely | Big Ideas in App Architecture
2026-06-17
Description
Most teams moving AI into production quickly discover that generating an output is the easy part; running it reliably, efficiently, and at scale is a discipline of its own. In this episode, David talks with Philip Kiely, engineer at Baseten and author of Inference Engineering, a free guide to building and operating AI inference systems. Philip argues that every company will soon need a dedicated team to own inference, and draws on his work at Baseten to explain why it demands fluency across GPU optimization, distributed systems, model correctness, and developer experience all at once. He walks through how agentic workloads are reshaping inference demands, which open-weight models are worth watching, and key optimization techniques including KV cache reuse, quantization, and speculative decoding. Grab your copy of Inference Engineering here: https://www.baseten.co/inference-engineering/