Search
Hot Posts
- A Comprehensive Guide to Multi-Node, Multi-GPU NVIDIA GPU Fabric Deployment
- Hands on DSH
- Dive into the Agent Loop
- How uv Works: Python Environments, Dependency Resolution, and Caching
- Cross-Attention in Multimodal AI: When One Stream Needs to Read Another
- LFT and Topology Exports in InfiniBand: Diagnose the Fabric Without Expanding the Attack Surface
- The Decision Engine Behind Fused Attention: How Transformer Engine Orchestrates cuDNN on Blackwell
- Megatron-Bridge in Practice: A Production 101 Guide to the Nemotron, Megatron-Core, and Transformer Engine Stack
- FMHA 101: From Transformer Attention to a CUDA Flash Attention Kernel
Recent Posts
- Hands on DSH
- Dive into the Agent Loop
- How uv Works: Python Environments, Dependency Resolution, and Caching
- Cross-Attention in Multimodal AI: When One Stream Needs to Read Another
- LFT and Topology Exports in InfiniBand: Diagnose the Fabric Without Expanding the Attack Surface
- The Decision Engine Behind Fused Attention: How Transformer Engine Orchestrates cuDNN on Blackwell
- Megatron-Bridge in Practice: A Production 101 Guide to the Nemotron, Megatron-Core, and Transformer Engine Stack
- FMHA 101: From Transformer Attention to a CUDA Flash Attention Kernel
- Enroot + Pyxis + SPANK: A Practical Architecture and Operations Guide for Slurm Containers
Tag Cloud
Linux (61)
XEN (29)
Life (27)
Memory (24)
Virtualization (23)
Diary (21)
C/C++ (19)
QEMU (17)
test (15)
CPU (14)
NVIDIA (13)
InfiniBand (11)
AI (11)
Interest (11)
Algorithm (11)
VisualStudio (11)
HPC (10)
GPU (10)
KVM (10)
Transformer (9)
eBPF (8)
MFC (8)
OpenClaw (7)
Debian (7)
LLM (6)
Engine (6)
MLOps (6)
Security (6)
TorchDynamo (6)
PyTorch (6)
CUDA (6)
programming (6)