FMHA 101: From Transformer Attention to a CUDA Flash Attention Kernel
A beginner-friendly, code-oriented introduction to Transformer attention, multi-head attention, online softmax, and a tiled CUDA FMHA forward kernel.
Read MoreA beginner-friendly, code-oriented introduction to Transformer attention, multi-head attention, online softmax, and a tiled CUDA FMHA forward kernel.
Read MoreDeep dive into eBPF VM implementation, Probe engine mechanisms, program counter jumps, and hardware resource access. Covers register architecture, instruction encoding, Kprobe/Kretprobe/Uprobe implementation, PMU access, JIT compilation optimization, and production best practices. Ideal for systems engineers and kernel developers.
Read More