Skip to content

Integrate vLLM 0.30.0 upgrade with current main - #425

Merged
jiangkuaixue123 merged 24 commits into
mainfrom
codex/v030-main-integration
Oct 9, 2026
Merged

jiangkuaixue123 merged 24 commits into
mainfrom
codex/v030-main-integration

Conversation

@jiaran-king

@jiaran-king jiaran-king commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose and scope

Integrate update/v0.30.0 with main through 6981ee1, incorporating #402,
#405, #407, #411, #412, #413, #417 and #424 while preserving main's changes,
including #427 and #359. Keep the pinned vLLM ced6857afa0ea7b2e3f0846a62e1394e90f15607
and vLLM-Ascend 8d4409d6256d8a6729140ddcc0d1889e3f96cdd6.

Integration changes

Known issues handled independently

Two fixes are extracted from this integration and are reviewed separately:

These are explicitly accepted known issues for this integration, not prerequisites
for merging #425. Each fix proceeds independently. No configuration-disabling
switch, compatibility framework or new runtime guard is added.

Qwen3.6 validation, new NPU MRV2 DBO, #422, performance qualification and inherited
DCO remediation remain outside scope.

CI image alignment

Use vllm/vllm-openai:v0.30.0 as the CI base. After installing dependencies,
report vLLM/PyTorch/CUDA versions and require vLLM 0.30.0.
Buildkite #151 stopped during image construction because the added uv pip check
rejected the official image's intentional NCCL 2.30.7 override of PyTorch's
2.29.7 pin. Remove that blanket check while retaining the version assertion.
Changed-file pre-commit hooks and embedded shell/Python syntax checks pass.
Image construction passed in #152; GPU configuration validation awaits the new Buildkite run.

GPU DBO configuration follow-up

Buildkite #152 passed image
construction, Simple Test and Legacy E2E. MRV2 reported three passes and three
failures: eager DP2 failed during startup rendezvous; eager/graph DBO timed out
on the first evaluation step.

a06d850 restores native MRV2 chunked-prefill and token-budget defaults, following
#416's configuration approach. Explicit caller choices, request concurrency,
sample counts, accuracy thresholds and live two-stage assertions are retained.
The candidate production stream-order change in 04ac606 is withdrawn;
afd_plugin/ and csrc/ match 3d30d17 exactly. Large-message P2P deadlock
investigation (#414) remains separate from this configuration validation.
Related CPU tests: 228 passed; applicable pre-commit hooks pass.
Fresh GPU CI is pending. The startup port race remains under observation.

Validation of the split

Integration head: a06d850fd9b5b580e172cafd518cce64d9d06612.

  • The four E2E runner/configuration/process-management unit files pass 222 tests.
  • Applicable pre-commit hooks pass for changed files, including Ruff, mypy,
    Markdown and SPDX; git diff --check passes.
  • A broader local selection reports 236 passed, 78 skipped, 15 failed.
    All 15 failures are missing-vLLM imports and reproduce on the original
    6941faf with the same selection. This is not a passing NPU runtime suite.
  • No GPU/NPU experiments were rerun during extraction. The hardware records below
    remain tied to their original source/configurations.
  • Recombining both extracted fixes restores afd_plugin/ and csrc/ exactly to
    6941faf. Test simplification and documentation are the remaining differences.

GPU hardware results

Hardware-tested source: 3ee2e812656a0d88baf6487b02fcf6f15a88ea2e; these are historical results; the restored MRV2 test defaults require fresh CI validation. H20-3e used pinned vLLM 0.30, PyTorch 2.13.0+cu130, CUDA 13.0 and NCCL 2.29.7. Target-runtime units: 1231 passed, 126 skipped, 165 subtests passed; skips require unavailable Ascend/NPU dependencies or devices. All 14 selected DeepSeek functional scenarios passed. E showed live two-stage eager execution; G showed eligible live two-stage FULL execution and matching FFN replay layouts.

Fixed first-300 GSM8K results (correct answers):

Model / path Strict match Flexible extract
DeepSeek-V2-Lite native 113/300 113/300
Legacy graph + DBO 2A2F 108/300 110/300
MRV2 B: eager, no DBO 116/300 118/300
MRV2 E: eager DBO 103/300 104/300
MRV2 G: FULL DBO 104/300 105/300
MRV2 B: single repeat 108/300 109/300
MRV2 E: single repeat 107/300 107/300
Qwen3-30B-A3B native 277/300 266/300
Qwen3 AFD graph 2A1F 282/300 273/300

All selected configurations meet the unchanged absolute threshold; that does not establish numerical equivalence. The initial B→E paired strict losses/gains were 15/2 and flexible 16/2; B→G were 13/1 and 14/1. In the one bounded B/E repeat, E−B changed from −13/−14 to −1/−2, with paired losses/gains 8/7 and 9/7. The original large gap did not consistently reproduce. Both rounds are retained, with no relaxed threshold or selection of the best run.

The repeat's two pytest/JUnit cases passed, but its controller raised at the immediate post-E GPU-process check. The E samples remain usable; verified final cleanup and allocation release were completed. The overall repeat job is not labelled PASS. Native/AFD topology and scheduling differences limit causal comparisons.

Existing old-main Legacy evidence is reused: same first256 questions and generation arguments, old vLLM 0.26 DBO 94/256 strict and flexible, integrated Legacy 93/256 strict, 94/256 flexible. Strict paired losses/gains were 8/7, flexible 8/8. Runtime/deployment differences remain; old main was not rerun. Old main does not support GPU MRV2 DBO E/G.

Qwen3 eager's existing128 result is 122/128 strict, 118/128 flexible; graph on the same128 is 123/128, 120/128. The completed DeepSeek native1319 result, obtained before the sample-scope reduction, was 495/1319 strict, 497/1319 flexible; its first300 are reused. Candidate300 results are representative comparisons, not full-dataset qualification. No further full-dataset, old-main or model-download experiments were run after the agreed scope reduction.

Historical NPU hardware results

The following runs included the CAMP padding/context repair that is now outside
this integration. They do not certify the split baseline without that repair,
and do not validate the later DSV2 shared-expert SP fix.

Pinned vLLM/Ascend, CANN 9.1.0, torch 2.10.0+cpu and torch-npu 2.10.0.post4 on Guian A3, using complete V2-Lite and DSV4 Flash W4A8 checkpoints. Five representative AFD configurations each ran the same first300 GSM8K questions, 8-shot, maximum generation512 and concurrency12. Flexible-extract results:

NPU configuration Correct Native reference Difference
V2-Lite MRV1 2A2F graph 106/300 106/300 0.00 pp
V2-Lite MRV2 2A2F FULL 113/300 106/300 +2.33 pp
V2-Lite ordinary Async CAM 106/300 106/300 0.00 pp
DSV4 W4A8 layered off 287/300 286/300 +0.33 pp
DSV4 W4A8 layered on 286/300 286/300 0.00 pp

Historical validation of the combined repairs in babcbd1 (including the MRV2 padding/context change now extracted from this PR): MRV2 2A1F eager/FULL_DECODE_ONLY/FULL each pass 25/25 after CAMP padding repair. The final common ABI shorts pass 75/75, all three exits0. Related target-environment CPU tests: 148 passed, no skips. Live traces confirm graph replay, gate/quantization, dispatch/combine and actual layered kernels; the workspace probe remains steady through896 work items.

#428 / #430 is fixed and integrated at 563164a. That source tree matches the hardware-tested #430 tip 2f506eedb4a4bd547368b8e8a9076b980d5c83c8. Native CAMP 72/72, Legacy DBO 2A1F 25/25, 2A2F 25/25, related units 24 passed, no skips. Live unequal stages were observed. Fixed first128 GSM8K: DBO on 41/128, off 40/128. These bounded repair checks supplement the five300 runs. The 2A2F model checks passed, but its shell wrapper later exited2; that orchestration error is disclosed in the repair evidence.

#429 is fixed at 2145e82. The original per-group operator case passed in three independent processes, and the full operator suite passed 16 tests, no skips, with the rebuilt kernel selected. Tolerances remain rtol=0.04, atol=0.05; the earlier approximately1.7% mismatch concerned output elements, not a fraction of operators. The five model-accuracy runs used per-channel weights. Failure, fix and device verification.

Shutdown AIV/MTE warnings occurred after scoring/SIGTERM; clean device exit is not claimed. Historical DeepSeek-V3.2 recipes have not been revalidated. The synchronous DSV4 A5/A3 profiles inherited from #359 check finished, nonempty concurrent responses, not answer accuracy; A5's existing content defect remains documented. Existing Qwen DBO exclusions also remain. These results do not qualify arbitrary topologies or establish performance equivalence.

Repository status

The split introduces one ordinary follow-up commit; it does not rewrite existing
contributors' history. The original fix commits remain in ancestry, but their
extracted changes are absent from this PR's final source tree.

The PR awaits final review. Current GitHub checks are shown on the PR;
old green checks are not attributed to this head. Historical DCO
identity mismatches were accepted through the authorized DCO override on
c900c87; checks on subsequent heads are reported separately. No temporary report, raw model
output or local review bundle is part of the committed changes.

yujuancao07 and others added 14 commits September 28, 2026 14:52
## Purpose

Upgrade the GPU backend target from vLLM `v0.26.0` to `v0.30.0`
(`ced6857a`). This re-bases every AFD GPU compatibility patch, worker /
model-runner hook, and model adapter onto the exact `v0.30.0` upstream
skeletons, refreshes the E2E harness for the new runtime contracts, and
updates all current-effectiveness documentation. PyTorch stays at
2.13.0;
the upgrade is a pure vLLM identity change on the GPU path.

## Issue

- Related issue(s): #389 
- Closing keyword, only if fully resolved: Closes #389 

## Scope

- In scope:
- GPU backend target `vllm==0.30.0` (`TARGET_VLLM_VERSION`, extras pin,
    `uv.lock`).
  - Rebase of the four GPU compat patches onto `v0.30.0` skeletons;
    removal of two workarounds absorbed upstream.
  - GPU workers / model runners (V1 + V2), FFN runner, ubatch wrapper.
  - Model adapters: DeepSeek-V2/V4, Qwen3 MoE, Qwen3.5/3.6.
  - E2E runner + fixtures; design docs, README, GPU guide, recipes.
- Out of scope:
  - NPU: remains tied to the vLLM `0.26.0` + vLLM-Ascend `80d8c194f`
    baseline and is **not revalidated for `0.30.0`** (follow-up work).
  - No changes to the vLLM source tree.


## Implementation Notes

- **Compat patches** (`afd_plugin/compat/patches/`): `run_engine_core`,
  `launch_core_engines`, `add_request_async`, `EngineCore.__init__/
_initialize_kv_caches/shutdown/run_busy_loop`, and `set_forward_context`
  re-copied from `ced6857a` with AFD marker blocks replayed. Two patch
  sets were **removed because upstream absorbed them** :
 the `DPCoordinatorProc` copies (0.30 gates
  wave/lockstep on `enable_wave_coordination`) and the async-Attention
request-count workaround (the base busy loop publishes counts natively).
  `config_validation` keeps the `deepep_low_latency` temp-backend bypass
  (0.30 validation point unchanged).
- **Workers/runners**: keyword-forward all drifted parameters
  (`max_num_sampled_tokens`, `randomize_inputs`, warmup/capture
  synchronize barrier, `full_cudagraph`/`pcp_manager`,
  `context_len`/`valid_dummy_state_slots`); mirror the upstream
`_max_full_descs_to_capture` truncation before building the AFD capture
  event tracker so `profile_only` memory profiling keeps event pairing;
  zero-return `profile_cudagraph_memory()` stub on the FFN runner; drop
  the `_create_sm_control_context` override (seam moved to
  `ubatch_utils.create_sm_control_context`). The `init_device`
  module-level runner swap seam was re-verified to still exist in 0.30.
- **Models**: `is_internal_router` replaced by the 0.30
  `MoERunner.gate is None` external-routing contract (Attention shells
  compute the gate before transport); `SparseMLAIndexGroupBuilder`
  plumbing and `is_fused_shared_expert_enabled` fields; DSV4 MoE now
  passes `num_hash_layers` with `use_sequence_parallel=False` pinned for
  AFD roles; Qwen3 MoE `is_fused_checkpoint_transposed` +
  `lm_head.tie_weights`; Qwen3.5/3.6 `replicate_shared_expert` /
  fused-shared-expert fields and `ParallelLMHead.tie_weights`.
- **E2E runner**: DBO decode threshold 1→2 (0.30 validates
  `dbo_*_token_threshold >= microbatch count`); new env knobs
  `AFD_E2E_API_PORT_BASE` (shared-host port isolation) and
`AFD_E2E_DBO_KEEP_CHUNKED_PREFILL=1` (0.30 requires chunked prefill for
  mamba cache mode `align`, i.e. hybrid Qwen3.5/3.6 DBO runs).
- Version gating consolidated into
`compat/vllm.is_target_vllm_compatible()`
  and applied to all patch install sites, including the previously
  ungated `engine_core` module.

## Test Plan

CPU-only (no GPU required):

- `uv run ruff check . && uv run ruff format --check .`
- `uv run pytest tests/unit -q` (full CPU-safe suite)

GPU-gated (8×H200, vLLM 0.30.0 + torch 2.13.0):

- Target-native control: vLLM 0.30.0 (plugins disabled) serves
  DeepSeek-V2-Lite, V2 model runner, eager, real request OK.
- Marker-based suites: `tests/e2e/models/deepseek_v2_lite` (all 13
  scenarios incl. V2 runner cells), `qwen3_moe`, `qwen3_6`.
- Full-dataset accuracy: `AFD_GSM8K_LIMIT=all` on
  `afd-eager-2a2f` / `afd-graph-2a2f`.

## Test Result

- CPU: full `tests/unit` on the final head against real vLLM 0.30.0 —
  0 failed 
- GPU DeepSeek-V2-Lite: **13/13 scenarios passed**
  (baseline-graph; afd-{eager,graph,graph-dbo}-2a2f;
afd-{eager,graph,graph-dbo}-2a1f; afd-v2-{eager,graph}-{1a1f,dp2,tp2}).
- GPU Qwen3-30B-A3B: **4/4** (baseline-graph;
  afd-{eager,graph,graph-dbo}-2a1f).
- GPU Qwen3.6-35B-A3B: **3/3** (baseline-graph; eager/graph-2a1f).
- Full GSM8K (1319 samples, threshold 0.27): afd-eager-2a2f
  strict-match 0.3821, afd-graph-2a2f 0.3829 vs native baseline 0.3723 /
  0.3760 — no accuracy regression.

## Docs Impact

- Files updated: `README.md`,

`docs/design/module/{index,execution_platforms,compatibility_and_patches,e2e_testing}.md`,
  `docs/gpu/NCCL_P2P_CONNECTOR_USER_GUIDE.md`,
  `docs/npu/CAM_P2P_CONNECTOR_USER_GUIDE.md`,
  `afd_plugin/connectors/README.md`,
  `recipe/README.md`,
  `recipe/gpu/P2pNcclAFDConnector/deepseek_v2_lite/README.md`.
- Historical 0.19.1-era recipes/notes intentionally preserved with their
  original labels; NPU support claims stay tied to the 0.26 pairing.

---------

Signed-off-by: yujuancao07 <yujuancao07@gmail.com>
## Purpose

Backport [#399](#399) to
`update/v0.30.0` so NPU custom operators can build with CANN 9.1 header
and logging changes.

## Issue

None. Source PR: #399.

## Scope

- In scope: NPU operator host include paths and logging headers.
- Out of scope: Operator computation and runtime scheduling.

## Implementation Notes

- Add the newer `include/op_common` header paths to host compilation and
`npu_op_code_gen` while retaining the existing `pkg_inc` paths.
- Replace `OpLogSub` with `D_OP_LOGE` in the operator error macro.
- Include `dlog_pub.h` before logging headers for the required module ID
and macro definitions.
- Cherry-pick the original commit with `-x`; the resulting patch ID
matches #399.

## Test Plan

Build and install the custom operators with CANN 9.1, then check
compatibility with an older CANN version.

## Test Result

- `git diff --check upstream/update/v0.30.0...HEAD` passed.
- Stable patch ID matches the commit from #399.
- Build and installation tests were not run because this environment has
no CANN toolkit.

## Docs Impact

- Files updated: None.
- If none, reason: Build compatibility fix with no user-facing API
change.

---

<details>
<summary>Essential PR Checklist</summary>

- [x] Purpose is clear and linked to public context when possible.
- [x] Scope is bounded.
- [x] Compatibility with vLLM v0.26.0 is considered; this backport
targets `update/v0.30.0`.
- [x] No changes are made to the vLLM source checkout.
- [x] Plugin-owned classes or explicit dotted class paths are preferred
over monkey patches; no classes or patches are added.
- [x] Any compat shim or monkey patch is isolated, idempotent,
version-guarded, documented, and tested; none is added.
- [x] Imports remain CPU-safe; no Python imports are changed.
- [x] Validation evidence is included, including skipped GPU tests when
applicable.
- [x] Documentation impact is stated.

</details>

Signed-off-by: lirx-pd <616517220@qq.com>
Co-authored-by: lirx-pd <616517220@qq.com>
## Purpose

Adapt the AFD NPU backend on `update/v0.30.0` to vLLM
`ced6857afa0ea7b2e3f0846a62e1394e90f15607` (0.30.0) and vLLM-Ascend
`8d4409d6256d8a6729140ddcc0d1889e3f96cdd6`, retaining Ascend
`ModelRunnerV1`. This follows the migration direction discussed in #344
while also covering Async CAM.

## Issue

- Related upgrade context: #344.

## Scope

- In scope: NPU config initialization and validation, target MRV1/device
metadata contracts, model-owned sequence parallelism, DeepSeek V2/V4
two-stage Async CAM, W8A8 force-load-balance routing, DBO stage
metadata, and focused unit/E2E coverage.
- Out of scope: switching to ModelRunnerV2 or changing upstream
vLLM/vLLM-Ascend source.

## Implementation Notes

- Replace old FlashComm1 assumptions with the target model's
`use_sequence_parallel_moe` layout, derived from EP/TP/all-to-all
configuration. Preserve Attention-side SP and stage-local token/Hash ID
alignment for Async CAM.
- Rebase the AFD compatibility patches onto the pinned Ascend revision;
keep plugin-owned model, runner, and connector behavior in their
existing modules.
- Review follow-up: consolidated adjacent config patch markers and
documented the unsupported Async CAM `mix_placement` path. The
force-balance patch removal experiment was reverted; that change will be
proposed separately.
- **Known graph limitation:** FULL graph replay corrupts accuracy for
live MLA DBO split steps. Such steps run eagerly while dummy capture and
unsplit graph steps retain FULL mode. This preserves correctness but
costs throughput on split steps. Remove the fallback after stage-local
MLA graph metadata can be updated safely.
- Changes are split into 36 signed-off commits, including dedicated
reverts of an unsuccessful DP-padding fix and the force-balance patch
removal experiment.

## Test Plan

- Compile/install the Ascend operators on 910C and run operator
numeric/communication checks.
- Run focused NPU unit tests and the DeepSeek V2/V4 Async CAM scenarios.
- Run the complete non-accuracy NPU marker suite and remaining full
accuracy cells before marking this PR ready for final merge.

## Test Result

- 910C operator editable build succeeded with `AFD_BUILD_ASCEND_OPS=1
SOC_VERSION=ascend910_9391 MAX_JOBS=8` (with
`SETUPTOOLS_SCM_PRETEND_VERSION_FOR_VLLM_AFD_PLUGIN=0.30.0` in the
synced container). Operator numerical checks and the 16-rank Async CAM
arithmetic/quantization checks passed.
- Focused target-runtime unit group: **114 passed**
(`test_npu_runtime.py`, `test_npu_ubatch_device.py`,
`test_deepseek_attention_metadata.py`, `test_camp2p_connector.py`). `git
diff --check` passed.
- DSV4 two-stage Async CAM on 16 NPU ranks: 10/10 arithmetic prompts
correct; scenario passed. DSV2 Attention-side SP DP2/TP2 plus FFN
DP4/EP4: HTTP inference smoke passed.
- Official DSV2 Lite graph-configured DBO, Attention DP2 + FFN DP2,
offline GSM8K **all 1,319 questions**: flexible exact match **0.3829**,
strict match **0.3798** (threshold 0.27); 26,560 live two-ubatch steps;
process exit 0. Used the existing container dataset cache without
downloading it. Live MLA split steps used the eager fallback described
above.
- After the review fixes, target-container config validation, Ascend
config lifecycle, and Async CAM connector unit selection: **59 passed, 1
skipped**. These review changes do not alter execution behavior.
- Remaining qualification: the complete non-accuracy marker suite and
all other full accuracy cells have not been recorded as passing. This PR
is a draft until those gates and documentation audit are complete.

## Docs Impact

- Updated `tests/e2e/README.md` for the target test configuration.
Public support/recipe documentation still needs a full target-version
audit before merge.

---

<details>
<summary>Essential PR Checklist</summary>

- [x] Purpose and scope are clear.
- [x] Existing vLLM v0.26.0 GPU compatibility was considered; changes
are NPU-scoped.
- [x] No changes were made to either upstream source checkout.
- [x] Compatibility patches are isolated and covered by focused tests.
- [x] Validation evidence and remaining gates are stated.
- [x] Documentation impact is stated.

</details>

---------

Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
## Summary

Restore the W8A8 force-load-balance patch removal from `2e702e4` as a
standalone change on top of the merged vLLM 0.30.0 NPU upgrade (#407).

- Remove the AFD patch of `AscendW8A8DynamicFusedMoEMethod` and its FFN
worker import.
- Stop accepting the retired
`additional_config.enable_force_load_balance` and
`force_load_balance_topn_per_rank` keys in the AFD Ascend config
wrapper.
- Document vLLM-Ascend's `additional_config.enable_force_eplb` as the
available FFN-side forced routing option.
- Keep the independent Attention-side Async CAM
`AFD_FORCE_BALANCED_TOPK_IDS` logic unchanged.

The native option does not provide the old `topn_per_rank` benchmark
mode and should not be treated as behaviorally identical. Historical
recipe files remain pinned to older release branches.

## Validation

- Restored diff has the same stable patch-id as `2e702e4`.
- `python3 tests/unit/compat/npu/test_ascend_config_lifecycle.py`: 5
tests run, 1 skipped, passed.
- `python3 -m compileall -q` on affected production Python files:
passed.
- `git diff --check`: passed.
- Read-only review found no blocking issues.

NPU E2E was not rerun for this standalone patch removal.

Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>

Signed-off-by: jiangkuaixue123 <jiangxiaozhou111@163.com>
## Purpose

Fix the NPU ModelRunner V2 wrappers left behind by the vLLM 0.30
upgrade. Native dummy/profile execution now passes additional state
arguments, and graph capture uses new input-preparation arguments and a
`profile_only` flag. The old
AFD signatures reject these calls. Graph-memory profiling also captures
only a prefix of FULL descriptors, while AFD previously expected events
for every descriptor.

This change aligns the wrappers and their expected capture events with
vLLM 0.30.0. Commit: `ab5b2c529ff8952e361b54e31e1494f917e54a4c`.

## Issue

- Related issue(s): #240 and #390 
- Closing keyword, only if fully resolved: None.

## Scope

- In scope: NPU V2 execution/capture interfaces, shared metadata type
declarations, and CPU regression tests for these contracts.
- Out of scope: Connector or kernel changes, vLLM source changes, and
full model/communication E2E qualification. The standalone workspace
test tools are outside this PR.

## Implementation Notes

- Forward `context_len` and `valid_dummy_state_slots` from
`AFDNPUAttentionModelRunnerV2.execute_model()` to the native runner.
- Forward `capture_model(profile_only=...)` and use the graph manager's
`_max_full_descs_to_capture` limit for AFD warmup/capture events.
- Match `prepare_inputs_to_capture()` with `full_cudagraph`,
`max_query_len`, and `pcp_manager`, replacing the obsolete `skip_attn`
argument.
- Declare the fields required by `AFDMetadataProviderMixin`, including
nullable pending metadata. Access the graph manager's instance
dictionary directly for temporary replay-hook state.

Native execution and capture remain delegated to vLLM-Ascend. The
existing capture hook restores the original input-preparation function
and AFD state on success and failure. These interfaces target the pinned
vLLM 0.30.0 contract; no runtime version guard is added.

## Test Plan

Run from the repository root. Use Python 3.10 or newer for pre-commit:

```bash
python -m pre_commit run --show-diff-on-failure --files \
  afd_plugin/v1/worker/attention_metadata.py \
  afd_plugin/v1/worker/npu/attention_model_runner_v2.py \
  tests/unit/v1/worker/test_npu_model_runner_v2.py
```

Run the CPU tests in the existing Ascend runtime environment. They
require the runtime packages but do not allocate an NPU or initialize a
connector:

```bash
VLLM_PLUGINS=ascend VLLM_USE_V2_MODEL_RUNNER=1 \
HF_HUB_OFFLINE=1 HF_DATASETS_OFFLINE=1 \
python -m pytest -o addopts= -q -p no:cacheprovider \
  tests/unit/v1/worker/test_npu_model_runner_v2.py \
  tests/unit/v1/worker/test_model_runner_v2.py \
  tests/unit/v1/worker/test_ffn_metadata.py \
  tests/unit/model_executor/models/test_forward_context.py
```

Full model startup, CAMP2P communication, logits comparison, and GSM8K
accuracy remain to be validated on the target A3 environment with local
model weights and dataset cache.

## Test Result

- Default pre-commit checks passed for all three changed files using
Python 3.10, including mypy and SPDX checks. No hooks were manually
skipped; hooks for unrelated file types reported no matching files.
- CPU regression suite: **119 passed**, including all **8** new NPU V2
contract cases. Coverage includes signatures, argument forwarding,
profiling descriptor limits, native dummy-run dispatch, and state
restoration after success or exceptions.
- A development snapshot passed a separate one-device check on **Ascend
910B1**: real NPU tensor operations and NPUGraph capture/replay, with
recorded AFD metadata and injected execution/replay failures. Native
runner execution, capture orchestration, input preparation, and
communication were stubbed. This run predates the final typing/SPDX
changes and is supporting component evidence, not final-commit E2E
validation.
- A3 connector and full model E2E tests were not run on this non-A3
machine. No model-accuracy or performance claim is made.

## Docs Impact

- Files updated: No user-facing documentation files. Patch comments now
identify the vLLM 0.30 interfaces and pinned source revision.
- If none, reason: This fixes internal runner compatibility without
adding configuration or changing the deployment workflow.

---

<details>
<summary>Essential PR Checklist</summary>

- [x] Purpose is clear and linked to public context when possible.
- [x] Scope is bounded.
- [x] Compatibility with vLLM v0.26.0 is considered.
- [x] No changes are made to the vLLM source checkout.
- [x] Plugin-owned classes or explicit dotted class paths are preferred
over monkey patches.
- [ ] Any compat shim or monkey patch is isolated, idempotent,
version-guarded, documented, and tested.
- [x] Imports remain CPU-safe; CUDA-heavy work is delayed or GPU-gated.
- [x] Validation evidence is included, including skipped GPU tests when
applicable.
- [x] Documentation impact is stated.

</details>

Signed-off-by: lirx-pd <616517220@qq.com>
…on (#412)

## Purpose

Fix expert selection, queued operator workspace lifetime, and MRV1 DP
stage coordination in the NPU v0.30 upgrade.

Follow-up to #407, based on `update/v0.30.0` after #411.

## Changes

- **Async CAM routing:** use Ascend's native router factory instead of
always constructing the grouped-top-k router. This restores native
fused/fallback selection and its expert IDs and weights, including
correction bias, normalization and scaling. Keep logical routed-only
IDs; remove unreachable shared-ID concatenation and reject unsupported
`mix_placement` at the helper boundary.
- **ACLNN workspace lifetime:** keep the workspace tensor alive in the
queued handler until kernel submission. Previously the tensor left scope
before its raw pointer was consumed; full-model Async CAM runs crashed.
The fix requires no additional device synchronization.
- **DP stage agreement:** include decode mode and microbatch eligibility
in the existing DP metadata collective so ranks make the same split
decision. Before native unsplit dummy execution, clear stale stage
slices and veto microbatching. This fixes an idle rank entering an extra
stage collective and hanging a single request after graph capture.

The production diff is limited to three files (+35/-32 lines), with
regression tests in two existing files.

## Validation

Hardware validation used four Ascend 910C NPUs, CANN 9.1.0 and full BF16
DeepSeek-V2-Lite, with vLLM `ced6857afa0ea7b2e3f0846a62e1394e90f15607`
and vLLM-Ascend `8d4409d6256d8a6729140ddcc0d1889e3f96cdd6`.

- Explicit small-tensor routing expectations cover four CPU fallback and
four real NPU fused cases. The previous router fails the NPU regression
cases.
- Runtime and ubatch tests: **102 passed**, including two-process Gloo
stage/collective checks and a stale-dummy-state regression that fails
before the fix.
- Routing, CAM registration and GMM tests after the workspace fix: **66
passed, 2 skipped**. The two opt-in GMM runtime cases were then enabled
separately: **2 passed**.
- Full-model Async CAM: **two consecutive successful runs**, with
correct generated arithmetic answers and normal asynchronous execution.
- Synchronous DP2 DBO: **13/13 requests returned correct answers**. Both
Attention ranks produced the same ordered sequence of 80 control
records, each containing 32 split and 48 unsplit records. Split
execution used the existing eager fallback.

These NPU results were collected on `78ea4ae` before rebasing onto #411.
All three modified production files are byte-identical after rebasing;
the test ASTs are unchanged (unrelated formatting edits were removed).
On the rebased tree, config lifecycle tests passed (5 run, 1 skipped),
Python compilation and `git diff --check` passed.

---------

Signed-off-by: zzh <jiaranran2@gmail.com>
…layer (#417)

## Purpose

Align the AFD attention-side gate precision and split-FFN activation
parameters with the native Ascend MoE path. Reuse one native router per
attention-side MoE layer.

Based on `upstream/update/v0.30.0` at `3352e75` (#412), which already
delegates CAM routing to the native factory and rejects `mix_placement`.
The routing change in this PR adds per-layer reuse to that upstream
implementation.

## Issue

- Related issue(s): None.
- Closing keyword: N/A.

## Scope

- In scope: Attention-side gate precision, per-layer native router
reuse, MoE activation parameters, and regression coverage.
- Out of scope: CAM transport changes, DBO scheduling, and performance
tuning.

## Implementation Notes

- Replace `ReplicatedLinear` with native `GateLinear` and
`_get_moe_router_dtype(config)`, preserving gate checkpoint names.
- Create one native Ascend router per attention-side MoE layer and reuse
it during forward execution. `select_cam_experts` also uses the native
factory when no cached router is supplied.
- Read SwiGLU limit, alpha, beta, and Situ parameters from
`experts.moe_config`, matching native defaults for unquantized, W8A8,
and W4A8 paths. Use the same config for the shared-expert clamp limit.
- Keep Ascend router construction in the NPU helper with deferred
imports. No new monkey patches or upstream source changes are
introduced.

### Why the router is reused per layer

The router belongs to the MoE layer and shares its lifetime with the
gate. This matches the native `FusedMoEFactory` construction path, where
the runner receives and retains `self.router`. It provides the same
ownership model for later MRV2 and native remote-MoE integration.
Inputs, top-k outputs, and CAM transfer state remain local to each call;
fused/fallback checks still run for the current input.

The current Ascend gate produces FP32 logits, its correction bias, when
present, is FP32, and checkpoint loading updates the bias in place.
These conditions preserve the router's bias reference. A future change
to logits dtype or replacement of the bias parameter must revisit that
reference. Reuse avoids repeated factory calls and Python object
construction, but no end-to-end performance gain is claimed.

Future native DBO integration must validate mutable router capture
callbacks and replay buffers across microbatches. Native remote-MoE
integration should transfer ownership to the native runner and remove
the temporary AFD `cam_router`/`create_gate_router` plumbing, leaving
one router per layer.

## Test Plan

CPU regression tests:

```bash
python -m pytest -m 'not npu' \
  tests/unit/model_executor/models \
  tests/unit/connectors/test_async_cam_connector.py \
  tests/unit/model_executor/test_deepseek_v4_hash_ids.py \
  tests/unit/model_executor/test_deepseek_v4_npu_weight_roles.py
```

NPU component tests, run from the repository root inside the Ascend
container:

```bash
npu run --gpus 1 --nonblock -- \
  env AFD_RUN_ASCEND_OP_RUNTIME=1 \
  python -m pytest -m npu \
    tests/npu/test_attention_gate_numerics.py \
    tests/unit/connectors/test_async_cam_connector.py
```

Run Python syntax checks, Ruff lint and formatting checks, the
repository mypy wrapper, SPDX checks, and `git diff --check`.

## Test Result

Validated commit: `b3372e4`.

- CPU regression: **328 passed, 4 deselected**, with NPU initialization
explicitly blocked by the test runner. The four NPU-only cases passed in
the separate device run below.
- NPU components on one Ascend 910B1: **18 passed, 36 deselected**. The
deselected cases are CPU tests covered by the CPU run.
- 2 gate tests covering FP32 checkpoint loading and logits with
FP16/BF16 inputs.
- 8 router comparisons against the native factory, covering scoring
functions, renormalization, and correction bias.
  - 4 fused-routing tests checking explicit expert IDs and weights.
- 4 unquantized expert arithmetic tests covering FP16/BF16 and
clamped/unclamped activation.
- Python syntax checks, Ruff lint and formatting checks, the repository
mypy wrapper, SPDX checks, and `git diff --check` all passed.

Tested with vLLM `ced6857a` (v0.30.0) and vLLM-Ascend `99d96c3d`.

## Docs Impact

- Files updated: None.
- Reason: Internal numerical alignment; no user-facing configuration or
API changes.

---

<details>
<summary>Essential PR Checklist</summary>

- [x] Purpose is clear and linked to public context when possible.
- [x] Scope is bounded.
- [x] No changes are made to the vLLM source checkout.
- [x] Plugin-owned classes or explicit dotted class paths are preferred
over monkey patches.
- [x] Any compat shim or monkey patch is isolated, idempotent,
version-guarded, documented, and tested. N/A: none added.
- [x] Imports remain CPU-safe; CUDA-heavy work is delayed or GPU-gated.
- [x] Validation evidence is included, including skipped GPU tests when
applicable.
- [x] Documentation impact is stated.

</details>

---------

Signed-off-by: lirx-pd <616517220@qq.com>
## Purpose

Enable GPU ModelRunnerV2 eager DBO and two-microbatch FULL CUDA Graph
replay on `update/v0.30.0`. Previously the GPU validator rejected DBO,
and native MRV2 stage preparation did not attach the AFD
metadata/control needed by the remote FFN.

## Issue

- Source implementation:
[jiaran-king/afd-mrv2-dbo-graph@71e6532](jiaran-king/afd-mrv2-dbo-graph@71e6532).
- No closing issue.

## Scope

- GPU MRV2, synchronous `P2pNcclAFDConnector`,
`compute_gate_on_attention=false`, two microbatches, Attention DP > 1.
- Three production files, focused regression tests, B/E/G E2E scenarios
and their ready/merge CI selections, and support documentation.
- No FFN execution rewrite, NPU DBO changes, dependency changes, or
modifications to installed vLLM. Historical handoff reports, raw traces,
and diagnostic controllers are not included.

## Implementation Notes

Wrap the native `UBatchRunner.prepare` on the current instance,
preserving native input slicing, Attention metadata and thread/stream
scheduling. Attach isolated per-stage AFD metadata and publish the
complete stage control before execution. Extend capture event matching
to include the microbatch count. For multi-stage FULL replay, reuse the
already-published control instead of sending the old single-stage
payload again. Restore temporary methods and state on scope exit.

The implementation targets vLLM
`ced6857afa0ea7b2e3f0846a62e1394e90f15607`. The three production files
preserve the reviewed implementation; compared with that snapshot, they
retain the newer upgrade branch's metadata type declarations/import and
use equivalent `manager.__dict__` access instead of `vars(manager)` so
mypy recognizes the mutable instance dictionary. Test-only type
annotations/casts and SPDX headers satisfy the repository hooks without
changing the execution protocol. Runtime evidence probes are test-only
and loaded through the native worker extension entry point. Acceptance
requires every expected A/F rank, live nonempty stages, consecutive FULL
executions, and matching FFN layouts within the evaluation window; it
does not claim exact cross-role transaction pairing.

## Test Plan

```bash
python -m pytest tests/unit/test_e2e_runner.py tests/unit/config/test_validation.py -q -o addopts=''
pre-commit run --files <changed-files>
```

GPU scenarios: `afd-v2-eager-dp2`, `afd-v2-eager-dbo-dp2`,
`afd-v2-graph-dbo-dp2`. The two DBO rows are also included in the
ready/merge Buildkite MRV2 selections.

## Test Result

- Local focused tests on this branch: **176 passed**; Ruff passed. Newly
added unit-test code is consolidated to a net 322 lines (132 fewer);
repeated scenario checks and redundant configuration combinations were
removed.
- All applicable pre-commit hooks passed, including Ruff/format, mypy
3.10, SPDX, Markdown, YAML and Buildkite checks.
- The runtime-dependent `test_model_runner_v2.py` was skipped locally
because PyTorch/vLLM are unavailable. This revision only simplifies its
artificial shared-context fixture; it does not claim a fresh runtime
pass.
- The full CPU suite was attempted locally but collection requires
PyTorch, which is not installed in the local development environment; it
is not reported as passed.
- Prior H20 validation of the source implementation: **340 runtime unit
tests passed**; B/E/G GSM8K **41/128, 40/128, 43/128**, with all-rank
execution assertions passing. Fixed base DeepSeek-V2-Lite, 2A2F DP2 TP1,
8-shot, concurrency 12, max_num_seqs 8, MBT 4096. Accuracy threshold
remains 0.27.
- [Prior hardware evidence and numerical
limits](https://github.com/jiaran-king/afd-mrv2-dbo-graph/blob/71e65320388163bcc5b380936b389ef5ea9a91bf/handoff/h20-results/README.md):
same-input sample-64 diagnosis explains a near-tied logits/argmax
difference; outputs are not bitwise equivalent. No speedup or full
GPU/model-matrix claim.
- Hardware tests have **not been rerun on this rebased PR**. Existing
evidence is explicitly attributed to the source implementation.

## Docs Impact

README now describes the narrow GPU MRV2 DBO support instead of the
blanket unsupported statement. E2E documentation records the common
B/E/G conditions and execution-evidence requirements.

## Checklist

- [x] Scope and exact target version are stated.
- [x] Native MRV2 execution is retained; no upstream source edits.
- [x] Temporary AFD hooks are scoped, documented, and restored.
- [x] CPU checks and prior hardware evidence are distinguished.
- [x] This PR targets `update/v0.30.0`; promotion to `main` is a
separate step.

---------

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Preserve main workspace warmup, Legacy chunked-prefill choices, native NPU optimizations, and layered FFN support. Resolve metadata declarations, use routed-expert EPLB ownership, and keep MRV2 128-sample comparisons above the smoke-only threshold. GPU runtime validation remains pending; NPU hardware validation is deferred.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Keep DBO and MLA tests isolated from the separately tested AscendConfig installer. Allow existing pidfd mocks on Python builds without those optional OS attributes. The original behavior assertions remain unchanged. Validation: 25 focused tests and Ruff passed.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Express existing model and control-plane lifecycle contracts, preserve staged metadata narrowing, and type existing test fixtures. The repository changed-file mypy hook now passes for 33 production and 50 test files. Ruff passes; local affected tests: 51 passed, 11 runtime skips, 6 subtests passed. Independent production review found no behavior changes.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Pad many-A eager inputs before preparation, preserve FFN DP context, and match CAMP rank strides in receive counts and graph keys. Release the queued workspace capture after successful ACLNN submission while retaining its pre-submission ownership.

Document five 300-question NPU accuracy runs, real operator paths, bounded repair regressions, and remaining Legacy 2A1F and shutdown limits.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Wait for the UB-to-UB DataCopy before Quant reads the SwiGLU result and
reuses its workspace. The copy runs on PIPE_V, and automatic synchronization
is disabled for these kernels.

Clarify the existing device-test comments while preserving both calls, the
FP32 reference, and rtol=0.04/atol=0.05.

Validated on 910C: the original per-group case passed in three independent
processes, and all 16 cases in test_async_cam_layered_w4a8.py passed. Confirmed
the updated kernel was loaded. Ruff, SPDX, and git diff --check passed.

Refs #429.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Fix unequal live-stage CAMP payloads in NPU MRV1 Legacy DBO 2A1F.

Create independent physical control counts and pad native send tensors inside the opaque runtime operation, preserving original query lengths and real DBO.

Guian A3: 2A1F and 2A2F content checks each 25/25; native wire probe 72/72; related units 24 passed. Frozen GSM8K 128: DBO on 41/128, off 40/128.

Refs #428.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Preserve synchronous DSV4 deployment profiles alongside MRV2 DBO evidence and pinned runtime arguments. Refresh representative hardware validation docs, CAMP rank mapping and patch inventory; shorten the test probe introduction.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
@jiaran-king
jiaran-king force-pushed the codex/v030-main-integration branch from 07e5dbe to df0a57c Compare October 9, 2026 01:19
Gather model-local tokens before the TP-sharded shared MLP and select the local rows after its native output reduction. Preserve replicated weights and non-SP execution. Add a CPU arithmetic regression with distinct rank inputs and correct the shared-weight and EngineCore documentation.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Move MRV2 input padding and FFN context ownership, plus DSV2 SP shared-expert computation, to independent review branches. Retain the workspace release and strided counts/graph keys required by Legacy DBO. Document the accepted known gaps and scope historical validation to its tested code.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
Use the v0.30.0 base image, report the installed runtime versions and fail the image build on a wrong vLLM version or inconsistent dependencies.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
@jiaran-king jiaran-king added the merge-test Used to trigger merge CI in PRs. label Oct 9, 2026
Remove the blanket pip consistency check, which rejects the official vLLM image's intentional NCCL override. Keep the vLLM version assertion and runtime version output.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
@jiaran-king jiaran-king added merge-test Used to trigger merge CI in PRs. and removed merge-test Used to trigger merge CI in PRs. labels Oct 9, 2026
Queue each send/receive pair before yielding to the other microbatch. The previous shared-compute-stream order could block a large stage-one send ahead of the stage-zero receive while FFN waited to return stage zero. Use native DBO stream/event synchronization to retain overlap, preserving NPU transfer ordering.

Validated changed-file pre-commit and function-level checks with simulated dependencies. GPU execution and graph replay require CI verification.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
@jiaran-king jiaran-king added merge-test Used to trigger merge CI in PRs. and removed merge-test Used to trigger merge CI in PRs. labels Oct 9, 2026
Withdraw the unqualified production stream-order change from 04ac606 while evaluating the test configuration, as requested. Restore MRV2's native chunked-prefill behavior and token budget instead of forcing chunking off with a 4096-token budget. Preserve explicit caller choices, concurrency, sample counts, accuracy gates and live DBO assertions.

Validated 228 related CPU tests, applicable pre-commit hooks and diff checks. Production source matches 3d30d17; GPU validation remains pending.

Signed-off-by: Zhou ziheng <jiaranran2@gmail.com>
@jiaran-king jiaran-king added merge-test Used to trigger merge CI in PRs. and removed merge-test Used to trigger merge CI in PRs. labels Oct 9, 2026
@jiangkuaixue123
jiangkuaixue123 merged commit b7561d7 into main Oct 9, 2026
3 checks passed
@jiangkuaixue123
jiangkuaixue123 deleted the codex/v030-main-integration branch October 9, 2026 07:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-test Used to trigger merge CI in PRs.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants