A model architecture alone does not explain how a user request becomes a response. Between the interface and the result, state must be represented, scheduled, moved, synchronized, stored, and checked. Deep Dive into DeepSeek, Volume II makes that execution path visible.
Inside Volume IISystems foundations, model serving, scheduling, and runtime stateDeepSeek-V4 runtime behavior and speculative decodingFlashMLA, DeepGEMM, TileKernels, and TileLang kernel surfacesTensor, pipeline, expert, sequence, and context parallelismDeepEP communication and fused expert-parallel executionExpert placement with EPLB and pipeline scheduling with DualPipe3FS storage and distributed data processing with smallpondPrecision, quantization, and numerical correctnessLong-context execution, system oracles, replay, and failure analysisNVIDIA kernel measurements, benchmarking, and evaluation protocolsThe book follows the ownership and movement of request state across serving runtimes, GPU kernels, distributed communication, and storage. It identifies which component owns each handoff, keeps performance reasoning attached to correctness, and distinguishes conclusions established by source inspection from conclusions established by measurement.
Written for machine-learning engineers, systems practitioners, researchers, and students, this volume builds on the model vocabulary established in Volume I while reintroducing the systems contracts it needs. No prior GPU programming or distributed-systems programming experience is required.
For readers who want to understand what happens between a model interface and a measured GPU result, this is a mechanism-first guide to the DeepSeek systems stack.