Part V: Distributed Inference and Serving
Chapter 22: Per-Node Inference Efficiency: A Prerequisite

Per-Node Inference Efficiency: A Prerequisite

Part V is about serving a model across a fleet: many replicas, a load balancer, autoscaling on queue depth, failover across machines. None of that distributed machinery can be sized, costed, or reasoned about until you know what one machine actually does when it serves a request, and that single-node question is the whole of this chapter. Everything in the pages that follow is scale-up: how a single accelerator holds a model, how quantization shrinks the weights that must fit in its memory, how the KV cache grows with every generated token and eventually decides how many concurrent requests one node can hold, how an attention kernel reads and writes that cache, and how a compiler fuses the operations into something the hardware runs efficiently. This is deliberately not the main event of the book, which leads with distribution and treats single-node efficiency as a labeled prerequisite. It is here because a distributed serving system multiplies per-node behavior across every replica it runs, so a fleet built on a node that wastes half its memory bandwidth is the same waste repeated a thousand times and billed a thousand times over. The chapter develops the unit economics of one node, from the latency and throughput a single accelerator delivers, through the model-compression and attention-and-cache techniques that move those numbers, to the closing translation that turns a measured per-node profile into a replica count and a monthly bill. Read it as the calibration step before the distribution begins: you cannot do fleet arithmetic until you know the cost of the unit you are multiplying.

Conceptual illustration for Chapter 22: Per-Node Inference Efficiency: A Prerequisite

"They asked the load balancer how many replicas the fleet needs. The load balancer asked one node how fast it serves a token. One node said it was busy moving the KV cache around and would get back to everyone shortly."

A GPU Profiling Its Own Memory Bandwidth

Chapter Overview

This chapter opens Part V, the part of the book about serving models across a fleet, and it opens it on a single machine. The earlier parts spread one training run across many accelerators; now the book turns from building models to serving them, and serving has a unit cost that distribution multiplies. The nine sections develop that unit cost end to end. They are scale-up throughout, by design and by label, because the distributed serving chapters that follow assume you already know what a node does with a request and only need to reason about how many nodes to run and how to route between them.

The sections fall into three movements. The first names the stakes: Section 22.1 shows why one node's efficiency determines fleet cost, establishing latency, throughput, and memory footprint as the per-node numbers that every later replica inherits. The second movement is the toolkit of model compression and serving optimization, the techniques that move those numbers on a single accelerator: Sections 22.2 through 22.4 shrink the model itself through quantization, pruning and sparsity, and knowledge distillation; Sections 22.5 and 22.6 attack the memory and compute of attention through the KV cache with paged attention and through FlashAttention and efficient attention kernels; and Sections 22.7 and 22.8 raise single-node utilization through continuous batching and speculative decoding and through compilation and kernel optimization. The third movement closes the loop: Section 22.9 takes the measured per-node profile and turns it into fleet sizing, translating tokens per second and memory per sequence into a replica count and a cost, which is exactly the input the distributed serving chapters consume.

Read in order, the nine sections take you from "fleet cost is per-node cost multiplied" to a worked path for moving every per-node number that matters: place the model in less memory with quantization, prune and distill it smaller, manage the KV cache so a node holds more concurrent sequences, run attention with a kernel that respects the memory hierarchy, keep the accelerator busy with continuous batching and speculative decoding, compile the graph so the hardware runs it well, and finally convert the improved profile into how many replicas the fleet needs and what it will cost. The argument is cumulative and it points outward: every technique here is justified by the fleet that will multiply it, and the chapter ends by handing that multiplied arithmetic to Part V.

Prerequisites

This chapter assumes the performance vocabulary the book has already built. From Chapter 3: Scalability and Performance Models you carry the roofline model and the distinction between compute-bound and memory-bound execution, because the central fact of inference, that generating one token at a time is bound by memory bandwidth rather than arithmetic, is a roofline statement, and most of this chapter is a campaign to move work off the memory-bound side of that roofline. From the large-model training chapters, Chapter 16: Model, Pipeline, and Sharded Parallelism and Chapter 19: Training Foundation Models at Scale, you carry the size of modern models, the transformer architecture, the attention mechanism, and the practical fact that weights and activations do not fit comfortably in one accelerator's memory, which is precisely why quantization, the KV cache, and paged attention matter. Beyond that the chapter assumes comfortable Python, familiarity with the transformer and its attention computation, and a working sense of what an accelerator's memory and compute resources are. No prior experience with inference serving or model compression is needed; Section 22.1 builds the per-node cost framing from the ground up before any specific technique appears.

Learning Objectives

Chapter Roadmap

What's Next?

This chapter is a single-node prologue, and it exists to be multiplied. Having measured what one machine costs to serve, the book now returns to its real subject and spreads that machine across a fleet. Chapter 23: Distributed Inference Systems opens the distributed half of Part V, and it asks the questions a single node cannot answer: how many replicas of the node you just profiled the fleet needs, how a batch-aware load balancer routes requests across them, how the system autoscales on GPU utilization and queue depth, how it loads large models and survives cold starts, and how it stays available when a replica fails. Every one of those decisions takes the per-node numbers from this chapter, the tokens per second, the memory per sequence, the concurrency a node can hold, as its input, and reasons about coordination on top of them. Read it next, and watch the chapter you just finished become a single term in a much larger product: where Chapter 22 asked what one machine costs to serve, Chapter 23 asks how to run a thousand of them as one service.

Bibliography & Further Reading

Quantization

Frantar, E., Ashkboos, S., Hoefler, T., Alistarh, D. "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers." arXiv:2210.17323, 2022. arxiv.org/abs/2210.17323

The one-shot post-training method that quantizes large language models to low bit-width with second-order error correction, a core reference for the post-training quantization of Section 22.2.

📄 Paper

Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., Han, S. "AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration." arXiv:2306.00978, 2023. arxiv.org/abs/2306.00978

The activation-aware scheme that protects the most salient weights during quantization, one of the production-grade methods surveyed in Section 22.2.

📄 Paper

Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., Han, S. "SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models." arXiv:2211.10438, ICML 2023. arxiv.org/abs/2211.10438

The method that migrates quantization difficulty from activations to weights so both can be quantized to eight bits, addressing the activation-outlier problem of Section 22.2.

📄 Paper

Dettmers, T., Lewis, M., Belkada, Y., Zettlemoyer, L. "LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale." arXiv:2208.07339, NeurIPS 2022. arxiv.org/abs/2208.07339

The mixed-precision decomposition that isolates outlier features so the bulk of a transformer runs in eight-bit integers without accuracy loss, foundational to the quantization discussion of Section 22.2.

📄 Paper

Pruning and Sparsity

Frantar, E., Alistarh, D. "SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot." arXiv:2301.00774, ICML 2023. arxiv.org/abs/2301.00774

The one-shot pruning method that removes a large fraction of weights from a billion-parameter model without retraining, the central reference for Section 22.3.

📄 Paper

Sun, M., Liu, Z., Bair, A., Kolter, J. Z. "A Simple and Effective Pruning Approach for Large Language Models (Wanda)." arXiv:2306.11695, ICLR 2024. arxiv.org/abs/2306.11695

The pruning criterion that combines weight magnitude with input activation norm to prune without weight updates, a lightweight counterpart to SparseGPT in Section 22.3.

📄 Paper

Knowledge Distillation

Hinton, G., Vinyals, O., Dean, J. "Distilling the Knowledge in a Neural Network." arXiv:1503.02531, NeurIPS 2014 Deep Learning Workshop. arxiv.org/abs/1503.02531

The paper that introduced training a small student to match a large teacher's softened outputs, the foundation of the distillation technique in Section 22.4.

📄 Paper

KV Cache and Attention

Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., Stoica, I. "Efficient Memory Management for Large Language Model Serving with PagedAttention." arXiv:2309.06180, SOSP 2023. arxiv.org/abs/2309.06180

The vLLM paper that manages the KV cache in non-contiguous pages to eliminate fragmentation and raise concurrency, the direct reference for the paged attention of Section 22.5.

📄 Paper

Dao, T., Fu, D. Y., Ermon, S., Rudra, A., Ré, C. "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness." arXiv:2205.14135, NeurIPS 2022. arxiv.org/abs/2205.14135

The IO-aware kernel that computes exact attention without materializing the full score matrix, the central method of Section 22.6.

📄 Paper

Dao, T. "FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning." arXiv:2307.08691, 2023. arxiv.org/abs/2307.08691

The follow-up that improves work partitioning and parallelism to push attention closer to peak hardware throughput, extending the kernel discussion of Section 22.6.

📄 Paper

Batching and Decoding

Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., Chun, B.-G. "Orca: A Distributed Serving System for Transformer-Based Generative Models." OSDI 2022. usenix.org/conference/osdi22/presentation/yu

The system that introduced iteration-level (continuous) batching for generative transformers, the throughput technique developed in Section 22.7.

📄 Paper

Leviathan, Y., Kalman, M., Matias, Y. "Fast Inference from Transformers via Speculative Decoding." arXiv:2211.17192, ICML 2023. arxiv.org/abs/2211.17192

The method that drafts multiple tokens with a small model and verifies them in parallel with the large one, the latency technique paired with continuous batching in Section 22.7.

📄 Paper