How vLLM Works

An animated, step-by-step tour of the engine that serves large language models fast: the KV cache, the memory problem, PagedAttention, continuous batching, and prefix caching. Press Play or step through each chapter.