Skip to content

Extend FP32 MoE numerics to low-precision recipes and registered activations - #77

Draft
zianglih wants to merge 3 commits into
radixark:miles-mainfrom
zianglih:agent/moe-fp32-low-precision
Draft

Extend FP32 MoE numerics to low-precision recipes and registered activations#77
zianglih wants to merge 3 commits into
radixark:miles-mainfrom
zianglih:agent/moe-fp32-low-precision

Conversation

@zianglih

@zianglih zianglih commented Aug 5, 2026

Copy link
Copy Markdown

What does this PR do ?

@HumansAnd

Generalizes the moe_activation_in_fp32 and moe_combine_in_fp32 options introduced by radixark/Megatron-LM#68 across supported Transformer Engine low-precision routed-MoE paths. It also replaces the SwiGLU-specific activation implementation with a reusable registration surface, with SwiGLU and squared ReLU as built-ins.

Changes

  • Extends both FP32 MoE options to the standard FP8 recipes (delayed, tensorwise, blockwise, and mxfp8) in E4M3 and HYBRID formats, plus the standard NVFP4 recipe.
  • Adds a registry for FP32 expert activations keyed by the canonical activation callable and gating mode. Each registration binds activation-specific model-config context and supplies FP32 forward/backward callbacks.
  • Keeps router-probability weighting, its gradient, and the expert-GEMM boundary casts in one shared autograd wrapper, with no activation-specific branches.
  • Registers SwiGLU and squared ReLU. Activation-in-FP32 supports the fused-permutation path used by squared-ReLU MoE models and intentionally supersedes the lower-precision weighted squared-ReLU fusion when enabled.
  • Canonicalizes the YAML squared-ReLU callable so CLI, provider, and YAML configurations resolve the same registration.
  • Enforces the supported execution contract: nonlegacy TEGroupedMLP routed experts; a registered activation/gating pair for activation-in-FP32; and unfused AlltoAll with FP32 router probabilities for combine-in-FP32.
  • Adds precision-matrix, configuration-contract, registry-extension, YAML-identity, and FP32/BF16 forward/backward tests.
  • Checks in a reproducible EP8 smoke module at tests.functional_tests.test_cases.common.moe_fp32_low_precision for the low-precision runtime matrix.

Scope

This is a model- and architecture-neutral Megatron-core feature: there are no Inkling-, Nemotron-, or other model-specific checks or call paths. No Miles, SGLang, or Transformer Engine source changes are required. The activation option covers routed TEGroupedMLP experts; shared experts remain outside this path. Custom and per-module quantization recipes remain outside the supported contract.

Validation

  • git diff --check
  • 68 targeted unit/config tests on bare image radixark/miles:dev-202608041247: all passed
  • Black 24.10.0 and isort checks on the fully owned helper/test modules; py_compile on all six changed Python files
  • 21/21 fresh 8x B200 EP8/TP1 runtime launches passed on PyTorch 2.11.0+cu130 and Transformer Engine 2.17.0:
    • all five recipe families (delayed, tensorwise, blockwise, mxfp8, and nvfp4) in SwiGLU activation+combine mode and combine-only mode
    • all five recipe families with registered squared ReLU, activation-in-FP32, fused permutation, and AllGather
    • all four FP8 recipes in E4M3 activation+combine mode, in addition to HYBRID coverage
    • squared-ReLU activation+combine spots for MXFP8 and NVFP4

The runtime harness verifies the active Transformer Engine recipe class, quantized FC1 and FC2 grouped GEMMs, one local expert per EP8 rank, cross-rank routing, quantization padding, FP32 router probabilities reaching the final unfused AlltoAll unpermute when combine-in-FP32 is enabled, the registered FP32 activation autograd node, and finite nonzero hidden/router/expert gradients over two iterations.

⚠️ For major changes (either in lines of code or in its impact), please make sure to first share a design doc with the team. If you're unsure what's the best way to do so, contact the @mcore-oncall.

Contribution process

flowchart LR
    A[Pre-checks] --> B[PR Tests]
    subgraph Code Review/Approval
        C1[Expert Review] --> C2[Final Review]
    end
    B --> C1
    C2 --> D[Merge]
Loading

Pre-checks

  • I want this PR in a versioned release and have added the appropriate Milestone (e.g., Core 0.8)
  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

The following process is enforced via the CODEOWNERS file for changes into megatron/core. For changes outside of megatron/core, it is up to the PR author whether or not to tag the Final Reviewer team.

For MRs into `main` branch

Feel free to message or comment the @mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

(Step 1): Add PR label Expert Review

(Step 2): Collect the expert reviewers reviews

  1. Attach the Expert Review label when your PR is ready for review.
  2. GitHub auto-assigns expert reviewers based on your changes. They will get notified and pick up your PR soon.

⚠️ Only proceed to the next step once all reviewers have approved, merge-conflict are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

(Step 3): Final Review

  1. Add Final Review label
  2. GitHub auto-assigns final reviewers based on your changes. They will get notified and pick up your PR soon.

(Optional Step 4): Cherry-pick into release branch

If this PR also needs to be merged into core_r* release branches, after this PR has been merged, select Cherry-pick to open a new PR into the release branch.

For MRs into `dev` branch The proposed review process for `dev` branch is under active discussion.

MRs are mergable after one approval by either eharper@nvidia.com or zijiey@nvidia.com.

Merging your PR

Any member of core-adlr and core-nemo will be able to merge your PR.

@zianglih zianglih changed the title [tml] extend fp32 MoE activation and combine to FP8 and NVFP4 [tml] extend fp32 MoE numerics to low precision and squared ReLU Aug 5, 2026
@zianglih zianglih changed the title [tml] extend fp32 MoE numerics to low precision and squared ReLU Extend FP32 MoE numerics to low-precision recipes and registered activations Aug 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant