lyan
CodeFun
  • Home
  • Archives
  • About
  • Links
EN 中文 日本語
  1. Home
  2. Archives

LFT and Topology Exports in InfiniBand: Diagnose the Fabric Without Expanding the Attack Surface

lyan 2026-06-26

A practical, evidence-based framework for collecting and handling InfiniBand LFT, topology, and UFM snapshot data without exposing the management plane.

Network NVIDIA RDMA InfiniBand HPC Security UFM
Read More

The Decision Engine Behind Fused Attention: How Transformer Engine Orchestrates cuDNN on Blackwell

lyan 2026-06-07

A source-grounded tour of Transformer Engine’s fused-attention dispatch: cuDNN Graph API, versioned backend eligibility, Blackwell paths, kernel-pipelining context, and deterministic training constraints.

GPU cuDNN Blackwell Transformer Engine Fused Attention Computing
Read More

Megatron-Bridge in Practice: A Production 101 Guide to the Nemotron, Megatron-Core, and Transformer Engine Stack

lyan 2026-05-25

A technical guide to the architecture, repository layout, parallelism-aware conversion mechanics, and an import→fine-tune→export workflow for NVIDIA’s Megatron stack.

NVIDIA Training MLOps Megatron Nemotron LLM
Read More

FMHA 101: From Transformer Attention to a CUDA Flash Attention Kernel

lyan 2026-05-05

A beginner-friendly, code-oriented introduction to Transformer attention, multi-head attention, online softmax, and a tiled CUDA FMHA forward kernel.

GPU programming CUDA FlashAttention Transformers Transformer Engine
Read More

Enroot + Pyxis + SPANK: A Practical Architecture and Operations Guide for Slurm Containers

lyan 2026-04-26

A practical guide to integrating Enroot, Pyxis, and Slurm SPANK for unprivileged GPU and MPI containers, from architecture and deployment to troubleshooting.

NVIDIA HPC Slurm Containers MLOps
Read More
  • Previous
  • 2
  • Next

Search

Top Posts

  • Step into the world of ARM Server
  • Running Debian ARM64 on QEMU with UEFI

Hot Posts

  • A Comprehensive Guide to Multi-Node, Multi-GPU NVIDIA GPU Fabric Deployment
  • Miles, Slime, and Ray: Orchestrating the Post-Training Loop
  • Hands on DSH
  • Dive into the Agent Loop
  • How uv Works: Python Environments, Dependency Resolution, and Caching
  • Cross-Attention in Multimodal AI: When One Stream Needs to Read Another
  • LFT and Topology Exports in InfiniBand: Diagnose the Fabric Without Expanding the Attack Surface
  • The Decision Engine Behind Fused Attention: How Transformer Engine Orchestrates cuDNN on Blackwell
  • Megatron-Bridge in Practice: A Production 101 Guide to the Nemotron, Megatron-Core, and Transformer Engine Stack

Recent Posts

  • Miles, Slime, and Ray: Orchestrating the Post-Training Loop
  • Hands on DSH
  • Dive into the Agent Loop
  • How uv Works: Python Environments, Dependency Resolution, and Caching
  • Cross-Attention in Multimodal AI: When One Stream Needs to Read Another
  • LFT and Topology Exports in InfiniBand: Diagnose the Fabric Without Expanding the Attack Surface
  • The Decision Engine Behind Fused Attention: How Transformer Engine Orchestrates cuDNN on Blackwell
  • Megatron-Bridge in Practice: A Production 101 Guide to the Nemotron, Megatron-Core, and Transformer Engine Stack
  • FMHA 101: From Transformer Attention to a CUDA Flash Attention Kernel

Tag Cloud

Linux (61) XEN (29) Life (27) Memory (24) Virtualization (23) Diary (21) C/C++ (19) QEMU (17) test (15) AI (14) CPU (14) NVIDIA (13) InfiniBand (11) Interest (11) Algorithm (11) VisualStudio (11) HPC (10) GPU (10) KVM (10) LLM (9) Transformer (9) eBPF (8) Agent (8) MFC (8) OpenClaw (7) Debian (7) Engine (6) Megatron (6) MLOps (6) Security (6) TorchDynamo (6) PyTorch (6)
lyan
CodeFun

©2026 xryan.net.

Code all fun things in the world by Leon.

  • Links
  • About Us
  • Feedback
QR Code
Scan to follow