lyan
CodeFun
  • Home
  • Archives
  • About
  • Links
EN 中文 日本語
  1. Home
  2. Archives
  3. Tag:Nemotron

Megatron-Bridge in Practice: A Production 101 Guide to the Nemotron, Megatron-Core, and Transformer Engine Stack

2026-08-26

A technical guide to the architecture, repository layout, parallelism-aware conversion mechanics, and an import→fine-tune→export workflow for NVIDIA’s Megatron stack.

NVIDIA MLOps Megatron LLM Training Nemotron
Read More

Search

Top Posts

  • Step into the world of ARM Server
  • Running Debian ARM64 on QEMU with UEFI

Hot Posts

  • A Comprehensive Guide to Multi-Node, Multi-GPU NVIDIA GPU Fabric Deployment
  • LFT and Topology Exports in InfiniBand: Diagnose the Fabric Without Expanding the Attack Surface
  • The Decision Engine Behind Fused Attention: How Transformer Engine Orchestrates cuDNN on Blackwell
  • Megatron-Bridge in Practice: A Production 101 Guide to the Nemotron, Megatron-Core, and Transformer Engine Stack
  • FMHA 101: From Transformer Attention to a CUDA Flash Attention Kernel
  • Enroot + Pyxis + SPANK: A Practical Architecture and Operations Guide for Slurm Containers
  • HPC-X, MPI, PMIx & NCCL: A Full Dissection of the GPU Cluster Communication Stack
  • Deep Dive into Server Memory Architecture: From DRAM Granules to NUMA Modes
  • UFM vs OpenSM: Understanding InfiniBand Fabric Management

Recent Posts

  • LFT and Topology Exports in InfiniBand: Diagnose the Fabric Without Expanding the Attack Surface
  • The Decision Engine Behind Fused Attention: How Transformer Engine Orchestrates cuDNN on Blackwell
  • Megatron-Bridge in Practice: A Production 101 Guide to the Nemotron, Megatron-Core, and Transformer Engine Stack
  • FMHA 101: From Transformer Attention to a CUDA Flash Attention Kernel
  • Enroot + Pyxis + SPANK: A Practical Architecture and Operations Guide for Slurm Containers
  • HPC-X, MPI, PMIx & NCCL: A Full Dissection of the GPU Cluster Communication Stack
  • Deep Dive into Server Memory Architecture: From DRAM Granules to NUMA Modes
  • UFM vs OpenSM: Understanding InfiniBand Fabric Management
  • The Great Interconnect Debate: NVLink, InfiniBand, UALink, and the Quest for AI Supremacy

Tag Cloud

Linux (61) XEN (29) Life (27) Memory (24) Virtualization (23) Diary (21) C/C++ (19) QEMU (17) test (15) CPU (14) InfiniBand (11) Interest (11) Algorithm (11) VisualStudio (11) HPC (10) NVIDIA (10) KVM (10) eBPF (8) AI (8) MFC (8) Debian (7) MLOps (6) TorchDynamo (6) PyTorch (6) CUDA (6) Kernel (6) OS (6) UALink (5) NVLink (5) RDMA (5) Jetson (5) OpenClaw (4)
lyan
CodeFun

©2026 xryan.net.

Code all fun things in the world by Leon.

  • Links
  • About Us
  • Feedback
QR Code
Scan to follow