GPU
Why Habu
A Triton bug, a one-line GEMM, and more than fifty lines I do not want to write.
GPU Kernels & Compilers: A Compressed Interview Course
From linear algebra to FlashAttention, MoE, and the XLA lowering pipeline — the working set for a GPU-kernel / ML-compiler interview, told with diagrams, the underlying math, and real kernels. Weighted toward the compiler / XLA / PTX axis.