Systems Performance Engineering — learning map
Eight books, 75 chapters. Click a chapter to read it. Inside a chapter, the arrows move to the previous and next chapter. The figures and diagrams from each chapter are placed in the text where they belong.
1. GPU architecture and kernels
Hardware overview: SMs, HBM, NVLink, PCIe
GPU architecture, CUDA basics, occupancy
Memory access patterns, coalescing, profiling
Occupancy tuning, warp efficiency, ILP
Kernel efficiency, arithmetic intensity, roofline
Intra-kernel pipelining, warp specialization, thread block clusters
Inter-kernel pipelining, streams, stream-ordered allocation
CUDA graphs, dynamic scheduling, device-side launch
CUTLASS and CuTe
Triton kernels
2. Distributed systems
Failure models, partial failure, Two Generals, FLP
Failure detection, heartbeats, phi-accrual
Leader election
CAP, consistency models, linearizability, CRDTs
Anti-entropy, gossip, Merkle trees
Distributed transactions, 2PC/3PC, consistent hashing
Consensus: Paxos, Raft, PBFT
Communication patterns, streaming
gRPC internals: HTTP/2 framing, encoding
Interceptors, deadlines, metadata, load balancing
gRPC in production
NCCL, collectives, InfiniBand, fabric tuning
GPU storage I/O, GPUDirect
3. Systems language (Rust)
Traits, generics, closures, iterators
Threads, channels, Send/Sync
Async and futures — language level
Unsafe code and raw pointers
FFI — calling into C and CUDA
Atomics
Memory ordering
Building a spin lock
Building channels
Building Arc
Cache lines, false sharing, the processor
OS primitives: futex, parking
Building mutexes and condvars
Concurrency and parallelism models
Event queues, syscalls, epoll/kqueue/IOCP
Fibers and green threads
Futures in Rust
Coroutines and async/await desugaring
Runtimes, wakers, reactor-executor
Self-referential structs and Pin
Building your own runtime
4. PyTorch internals
Profiling, tuning, and scaling PyTorch
torch.compile, Inductor, Triton, XLA backends
5. LLM inference systems
Multinode inference, parallelism, decoding, routing
Speculative decoding
Profiling and debugging inference at scale
Disaggregated prefill and decode
KV-cache tuning, paged attention
Dynamic and adaptive engine optimizations
6. Quantization and compilers
FP8, NVFP4, GPTQ, AWQ, SmoothQuant, KV-cache quantization
LLVM IR syntax and IRBuilder
LLVM build system and tooling
LIT testing
TableGen
Clang architecture and AST
PassManager and AnalysisManager
Processing LLVM IR, writing passes
IR instrumentation and PGO
7. Infrastructure and operations
OS, Docker, Kubernetes tuning for GPU nodes
Scaling to very large GPU clusters
8. Reliability and benchmark practice
Profiling and tuning at scale
Optimization checklist (175+ items)
Suggested reading order
1. Fregly 6-9 -> write Triton kernels 2. Fregly 13-14 -> PyTorch and compiler 3. Fregly 15-19 -> inference, then read the vLLM source 4. Blandy 1-15 -> Rust basics 5. Blandy 19, 22 + Bos 1-6 -> concurrency 6. Samson 1-10 -> async runtime, then build the proxy project 7. Fregly 2-5 -> hardware and fabric 8. Fregly 10-12 -> advanced CUDA 9. Petrov 8-14 and Indrasiri 1-5 -> distributed theory and RPC 10. Pandey 1-3, then Hsu -> compilers, only if you target that work 11. Fregly 3, 20 -> infrastructure
Not in these eight books
AMD ROCm / HIP
no book — use AMD docs
MLIR
no book — use MLIR docs + Triton source
Kubernetes operators, CRDs, Terraform
no book — lowest priority
Books.
Fregly — AI Systems Performance Engineering — Chris Fregly
Petrov — Database Internals — Alex Petrov
gRPC — gRPC: Up and Running — Kasun Indrasiri & Danesh Kuruppu
Rust — Programming Rust, 2nd ed. — Blandy, Orendorff, Tindall
Rust Atomics — Rust Atomics and Locks — Mara Bos
Async Rust — Asynchronous Programming in Rust — Carl Fredrik Samson
LLVM Essentials — LLVM Essentials — Sarda & Pandey
LLVM Hsu — LLVM Techniques, Tips, and Best Practices — Min-Yih Hsu
Chapter text and figures were extracted from the PDFs. Running heads, page numbers, and figure whitespace were removed. Code listings keep their monospace layout. Nothing was reworded.