Systems Performance Engineering — learning map

Eight books, 75 chapters. Click a chapter to read it. Inside a chapter, the arrows move to the previous and next chapter. The figures and diagrams from each chapter are placed in the text where they belong.

1. GPU architecture and kernels

Hardware overview: SMs, HBM, NVLink, PCIe
GPU architecture, CUDA basics, occupancy
Memory access patterns, coalescing, profiling
Occupancy tuning, warp efficiency, ILP
Kernel efficiency, arithmetic intensity, roofline
Intra-kernel pipelining, warp specialization, thread block clusters
Inter-kernel pipelining, streams, stream-ordered allocation
CUDA graphs, dynamic scheduling, device-side launch
CUTLASS and CuTe
Triton kernels

2. Distributed systems

Failure models, partial failure, Two Generals, FLP
Failure detection, heartbeats, phi-accrual
Leader election
CAP, consistency models, linearizability, CRDTs
Anti-entropy, gossip, Merkle trees
Distributed transactions, 2PC/3PC, consistent hashing
Consensus: Paxos, Raft, PBFT
gRPC fundamentals and protobuf
Communication patterns, streaming
gRPC internals: HTTP/2 framing, encoding
Interceptors, deadlines, metadata, load balancing
gRPC in production
NCCL, collectives, InfiniBand, fabric tuning
GPU storage I/O, GPUDirect

3. Systems language (Rust)

Ownership, moves, references, lifetimes
Traits, generics, closures, iterators
Error handling, crates, modules
Threads, channels, Send/Sync
Async and futures — language level
Unsafe code and raw pointers
FFI — calling into C and CUDA
Atomics
Memory ordering
Building a spin lock
Building channels
Building Arc
Cache lines, false sharing, the processor
OS primitives: futex, parking
Building mutexes and condvars
Concurrency and parallelism models
Event queues, syscalls, epoll/kqueue/IOCP
Fibers and green threads
Futures in Rust
Coroutines and async/await desugaring
Runtimes, wakers, reactor-executor
Self-referential structs and Pin
Building your own runtime

4. PyTorch internals

Profiling, tuning, and scaling PyTorch
torch.compile, Inductor, Triton, XLA backends

5. LLM inference systems

Multinode inference, parallelism, decoding, routing
Speculative decoding
Profiling and debugging inference at scale
Disaggregated prefill and decode
KV-cache tuning, paged attention
Dynamic and adaptive engine optimizations

6. Quantization and compilers

FP8, NVFP4, GPTQ, AWQ, SmoothQuant, KV-cache quantization
LLVM IR syntax and IRBuilder
LLVM build system and tooling
LIT testing
TableGen
Clang architecture and AST
PassManager and AnalysisManager
Processing LLVM IR, writing passes
IR instrumentation and PGO

7. Infrastructure and operations

OS, Docker, Kubernetes tuning for GPU nodes
Scaling to very large GPU clusters

8. Reliability and benchmark practice

Profiling and tuning at scale
Optimization checklist (175+ items)

Suggested reading order

1. Fregly 6-9 -> write Triton kernels
2. Fregly 13-14 -> PyTorch and compiler
3. Fregly 15-19 -> inference, then read the vLLM source
4. Blandy 1-15 -> Rust basics
5. Blandy 19, 22 + Bos 1-6 -> concurrency
6. Samson 1-10 -> async runtime, then build the proxy project
7. Fregly 2-5 -> hardware and fabric
8. Fregly 10-12 -> advanced CUDA
9. Petrov 8-14 and Indrasiri 1-5 -> distributed theory and RPC
10. Pandey 1-3, then Hsu -> compilers, only if you target that work
11. Fregly 3, 20 -> infrastructure

Not in these eight books

AMD ROCm / HIP
no book — use AMD docs
MLIR
no book — use MLIR docs + Triton source
Kubernetes operators, CRDs, Terraform
no book — lowest priority
Books.
Fregly — AI Systems Performance Engineering — Chris Fregly
Petrov — Database Internals — Alex Petrov
gRPC — gRPC: Up and Running — Kasun Indrasiri & Danesh Kuruppu
Rust — Programming Rust, 2nd ed. — Blandy, Orendorff, Tindall
Rust Atomics — Rust Atomics and Locks — Mara Bos
Async Rust — Asynchronous Programming in Rust — Carl Fredrik Samson
LLVM Essentials — LLVM Essentials — Sarda & Pandey
LLVM Hsu — LLVM Techniques, Tips, and Best Practices — Min-Yih Hsu
Chapter text and figures were extracted from the PDFs. Running heads, page numbers, and figure whitespace were removed. Code listings keep their monospace layout. Nothing was reworded.