Skip to content

[WIP] Add NVFP4 fake QAT for grouped experts - #75

Draft
zianglih wants to merge 2 commits into
radixark:miles-mainfrom
zianglih:agent/nvfp4-qat
Draft

[WIP] Add NVFP4 fake QAT for grouped experts#75
zianglih wants to merge 2 commits into
radixark:miles-mainfrom
zianglih:agent/nvfp4-qat

Conversation

@zianglih

@zianglih zianglih commented Jul 31, 2026

Copy link
Copy Markdown

What does this PR do ?

Adds opt-in Transformer Engine-backed NVFP4 fake quantization-aware training for routed MoE grouped FC1/FC2 weights.

Warning

Experimental / WIP. This is the Megatron half of the NVFP4 QAT RL path and has not yet been validated on a GPU devbox.

Implementation

This mirrors the existing INT4 fake-QAT path in TEGroupedLinear._get_weight_tensors:

  • OPEN_TRAINING_NVFP4_FAKE_QAT_FLAG=1 creates a cached public te.pytorch.NVFP4Quantizer.
  • Each routed-expert FC1/FC2 weight is quantized and dequantized during forward while the original parameter remains the BF16 training master weight.
  • Transformer Engine's quantize and dequantize autograd functions provide straight-through gradients; the existing main_grad contract is retained.
  • The quantizer uses row-wise 1x16 NVFP4 weight scaling and pads only the row dimension required by Transformer Engine.

The quantizer mirrors Miles' existing NVFP4 conversion/export contract:

  • NVTE_NVFP4_4OVER6=weights|all enables weight-side 4over6.
  • NVTE_NVFP4_4OVER6_E4M3_USE_256=weights|all selects E4M3 max 256 when weight-side 4over6 is active; otherwise the bound is 448.
  • NVTE_NVFP4_4OVER6_ERR_MODE selects the 4over6 error metric and defaults to MAE.
  • Transformer Engine continues to consume its fast-math environment controls directly.
  • NVFP4 group size remains fixed at 16; no separate group-size environment variable is introduced.

Scope and limitations

This PR affects routed MoE expert FC1/FC2 weights using TE grouped linear only. It does not add fake quantization to dense MLPs, shared experts, sequential or legacy experts, activations, optimizer state, parameter gather, or native FP4 GEMMs.

The current Miles Qwen recipe uses expert tensor parallel size 1. With expert TP greater than 1, the local-shard amax used here can differ from full-weight rollout conversion because the mirrored quantizer contract disables amax reduction. Generic expert-TP parity is therefore follow-up validation/work rather than a claim of this draft.

Dependencies

Validation

Passed static check:

  • git diff origin/miles-main...HEAD --check

The repository pre-commit hooks were attempted, but the pre-existing INT4 block in this file triggers formatting and missing-docstring changes. Those unrelated changes were deliberately excluded to keep this PR additive and NVFP4-only. Per request, no unit or functional tests, benchmarks, accuracy runs, or GPU/devbox validation were performed before opening this draft.

⚠️ For major changes (either in lines of code or in its impact), please make sure to first share a design doc with the team. If you're unsure what's the best way to do so, contact the @mcore-oncall.

Contribution process

flowchart LR
    A[Pre-checks] --> B[PR Tests]
    subgraph Code Review/Approval
        C1[Expert Review] --> C2[Final Review]
    end
    B --> C1
    C2 --> D[Merge]
Loading

Pre-checks

  • I want this PR in a versioned release and have added the appropriate Milestone (e.g., Core 0.8)
  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

The following process is enforced via the CODEOWNERS file for changes into megatron/core. For changes outside of megatron/core, it is up to the PR author whether or not to tag the Final Reviewer team.

For MRs into `main` branch

Feel free to message or comment the @mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

(Step 1): Add PR label Expert Review

(Step 2): Collect the expert reviewers reviews

  1. Attach the Expert Review label when your PR is ready for review.
  2. GitHub auto-assigns expert reviewers based on your changes. They will get notified and pick up your PR soon.

⚠️ Only proceed to the next step once all reviewers have approved, merge-conflict are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

(Step 3): Final Review

  1. Add Final Review label
  2. GitHub auto-assigns final reviewers based on your changes. They will get notified and pick up your PR soon.

(Optional Step 4): Cherry-pick into release branch

If this PR also needs to be merged into core_r* release branches, after this PR has been merged, select Cherry-pick to open a new PR into the release branch.

For MRs into `dev` branch The proposed review process for `dev` branch is under active discussion.

MRs are mergable after one approval by either eharper@nvidia.com or zijiey@nvidia.com.

Merging your PR

Any member of core-adlr and core-nemo will be able to merge your PR.

@hhaAndroid

Copy link
Copy Markdown

@zianglih Hello, Since this is targeting B-series cards, why is NVFP4 QAT still needed? I noticed you also support NVFP4 training for RL. If we directly use real FP4 computation, wouldn’t the training-inference consistency be better? Why do we need to support both paths?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants