Generating a few seconds of video can easily take a minute of compute. Closing that gap isn't AI - it's engineering. A layered stack of low-level optimizations, each with real tradeoffs, and several that will surprise you even if you've been doing systems work for years. This talk opens the black box of video inference optimization: attention kernels, quantization, compilation, CUDA graphs - what they actually buy you, what they silently cost you, and where the hardest problems come from places nobody warned you about. Based on hands-on experience optimizing LTX Video's inference engine.

VP Engineering, Inference @ Lightricks