Skip to content

DL Notes

A focused knowledge base for modern deep learning systems, from hardware fundamentals to LLM serving and neural rendering.

Start Here

  1. Hardware - GPU/NPU basics, memory systems, and practical GPU configuration
  2. Tensor Operations - core tensor manipulations, activation functions, and CUDA Graph notes
  3. AI Infra Overview - the layered structure of an end-to-end inference system
  4. Inference Request Lifecycle - how one request moves through the whole online system
  5. AI Infra Metrics - TTFT, TPOT, and TPS
  6. Roofline Model for Inference - compute ceilings, bandwidth ceilings, and how optimizations move bottlenecks
  7. Data Movement and Communication - memory hierarchy, runtime state flow, and cross-device tensor exchange
  8. KV Cache - KV memory semantics, paged blocks, and prefix reuse
  9. Serving Runtime - chunked prefill, admission control, and runtime-side stability
  10. Parallelism - DP and TP from the perspective of memory, throughput, and communication
  11. Decoding and Sampling - temperature, top-k, top-p, and min-p
  12. Speculative Decoding - verification, accepted length, and per-token latency
  13. Training Objective - autoregressive pre-training objective
  14. Position Encoding - RoPE, M-RoPE, and TM-RoPE from sequence to multimodal space-time
  15. Models - model-specific notes (Qwen3-Omni, DFlash) and practical serving commands
  16. Neural Graphics - NeRF and Flow Matching foundations

Documentation Map

Systems and Infrastructure

AI Infra

Models

Graphics and Generative Modeling

Scope

This site emphasizes practical understanding:

  • concise theory with equations where useful
  • implementation-minded notes and runnable snippets
  • serving and performance considerations for real workloads