Skip to content

[Bug]: NPU MRV1 2A1F Legacy DBO corrupts outputs with unequal live stage lengths #428

Description

@jiaran-king

2026-10-09 integration update: PR #430 preserves real Legacy DBO and passes the original 2A1F counterexample and 2A2F regression, 25/25 each. Frozen GSM8K 128-question accuracy: DBO on 41/128 (32.03%), off 40/128 (31.25%). See the verified diagnosis and evidence below in the follow-up comment; the original reproduction/history remain available.

Summary

During #425 NPU acceptance, synchronous MRV1 Legacy DBO on 2A1F returned HTTP
200 but generated corrupt content when real concurrent requests produced
unequal per-stage Attention token counts. The historical-recipe repeat passed
only 12/25 content checks: the serial request passed, but 13 of the 24
concurrent responses were garbled. DBO-disabled 2A1F eager/graph and Legacy
DBO 2A2F passed the same 25-response probe.

This tracks an existing Legacy NPU path. It is separate from new NPU MRV2 DBO
support and does not require taking over #422. Repair work resumed after the bounded Goal was approved. A verified fix is
now published in PR #430,
merged into #425's integration branch as
563164a.
The implementation and recorded device evidence have been reviewed against
current #425 head df0a57c. The reported bug is fixed and bug-level validation
is complete. Main integration remains tracked in #425, which is still open
and Draft; the fix has not yet merged into main.

Current environment

  • Related integration: Integrate vLLM 0.30.0 upgrade with current main #425. Initial hardware-tested source:
    3ee2e812656a0d88baf6487b02fcf6f15a88ea2e; the affected Legacy runner is
    retained in follow-up 605ba33ac27a4fc667d5d4ef93732d9c5a4f626c.
  • vLLM 0.30.0: ced6857afa0ea7b2e3f0846a62e1394e90f15607.
  • vLLM-Ascend: 8d4409d6256d8a6729140ddcc0d1889e3f96cdd6.
  • Ascend 910C/A3; openEuler 24.03 LTS-SP4; Python 3.12.13;
    CANN 9.1.0; torch 2.10.0+cpu; torch-npu 2.10.0.post4; driver 25.5.1.1.
  • Complete DeepSeek-V2-Lite BF16 checkpoint, 27 layers and 64 routed experts.
  • Source-installed AFD with packaged Ascend operators; upstream source trees
    are unchanged.

Reproduction

Use the existing NPU afd-graph-dbo-2a1f scenario with Attention DP2/TP1,
one FFN rank, block size 128, max model length 4096, max batched tokens 4096,
max sequences 8, prefix caching and chunked prefill disabled, and DBO
decode/prefill thresholds 2/8.

The acceptance workload is one serial completion followed by 24 completions
at concurrency 12. Four prompt lengths are 305, 369, 433, 497 tokens.
The following driver reuses the repository launcher and its existing accounting
smoke prompt; save it as reproduce_legacy_2a1f.py in the repository root:

from concurrent.futures import ThreadPoolExecutor
import json
import urllib.request
from tests.e2e import runner

def probe(args):
    def request(index):
        prompt = runner.ACCOUNTING_PROMPT.replace(
            "You are a professional accountant.",
            "You are a professional accountant." + "\n" * (index % 4) * 64,
        )
        body = dict(model=runner.served_model_name(args, "attention"),
                    prompt=prompt, temperature=0, max_tokens=64)
        req = urllib.request.Request(
            f"http://{args.api_host}:{runner.attention_api_port(args)}/v1/completions",
            data=json.dumps(body).encode(),
            headers={"Content-Type": "application/json"},
        )
        with urllib.request.urlopen(req, timeout=240) as response:
            return json.load(response)
    outputs = [request(0)]
    with ThreadPoolExecutor(max_workers=12) as pool:
        outputs.extend(pool.map(request, range(24)))
    with open("legacy-2a1f-responses.json", "w") as saved:
        json.dump(outputs, saved, indent=2)
    for response in outputs:
        text = response["choices"][0]["text"].strip()
        assert ("62.50" in text and "2.17" in text) or text in {"B", "B.", "B)"}, text

runner.run_gsm8k_evaluation = probe
raise SystemExit(runner.main())
python reproduce_legacy_2a1f.py \
  --model /path/to/DeepSeek-V2-Lite \
  --scenario afd-graph-dbo-2a1f --device-backend npu \
  --attention-devices 0,1 --ffn-devices 2 \
  --vllm-bin /path/to/venv/bin/vllm \
  --common-vllm-arg=--max-model-len --common-vllm-arg=4096 \
  --common-vllm-arg=--max-num-batched-tokens --common-vllm-arg=4096 \
  --common-vllm-arg=--max-num-seqs --common-vllm-arg=8 \
  --common-vllm-arg=--block-size --common-vllm-arg=128 \
  --common-vllm-arg=--gpu-memory-utilization --common-vllm-arg=0.75 \
  --common-vllm-arg=--no-enable-prefix-caching \
  --common-vllm-arg=--no-enable-chunked-prefill

The device evidence was collected using this workload through an acceptance
wrapper around the same launcher. The standalone packaging above has not been
rerun after repair work was paused. Check live stage records, not just DBO
configuration or capture records.

Actual behavior and confirmed evidence

  • All 25 responses were saved. Concurrent failures include repeated !,
    zero-width characters and unrelated multilingual tokens. All four prompt
    lengths have failures; this is not a GSM8K extraction or strict-match issue.
  • Whole-batch DP counts were already padded to [1668, 1668].
  • Live request-boundary splitting subsequently produced stage 0
    [802, 433] and stage 1 [866, 674]. Decode also produced [3, 2]
    and [4, 3].
  • The failing repeat contains 242 live split control records with
    is_graph_capturing=False, is_warmup=False, is_graph_replaying=False.
    The existing MLA split fallback is already active: corruption occurs during
    eager split execution.
  • CAMP A2E/E2A kernels use equal per-Attention strides. For stage 0, FFN
    receives a total count of 1235 and divides it into 617/617, while the A
    payloads contain 802/433 rows. Both row coverage and payload offsets disagree.

The equal-chunk communication contract is demonstrably violated. Explaining
the observed corruption by this mismatch is a source-based diagnosis, not a
claim that a unique root cause or successful repair has been established.

An isolated connector stage-padding prototype passed 20 related unit cases
but still passed only 10/25 device content checks. It is uncommitted and
unpublished. A subsequent observational probe was interrupted at the
requester's instruction and has no completion verdict. These results must not
be presented as a fix or as evidence that metadata padding alone is sufficient.

History and source pointers

Expected behavior and bounded follow-up

Existing Legacy 2A1F DBO should return correct contents for unequal real
workloads. A repair must reconcile per-stage physical send/receive sizes with
the kernel protocol while preserving MLA's true query lengths. Changing only
control counts, padding only the whole batch, or reapplying the existing graph
fallback does not establish a solution.

Acceptance should rerun this concurrent counterexample, demonstrate real
two-stage execution and correct physical sizes, and retain a Legacy 2A2F
regression check. Use a bounded same-question accuracy comparison if the
numerical path changes; no full 1319-question rerun or strict-match gate is
required. If a fallback is used instead, disclose the lost DBO overlap rather
than labelling it preserved DBO support.

Before submitting

Commit identity was corrected on 2026-10-09 to jiaran-king (Zhou ziheng). Integration commit 563164a replaces 0366406 with an identical source tree; hardware results and the archived source identities are unchanged.

Activity

  1. jiaran-king commented on Oct 8, 2026

    @jiaran-king
    CollaboratorAuthor

    The repair from PR #430 has been merged into #425's integration branch as 563164a. The merged tree is identical to the reviewed and hardware-validated #430 tip 2f506eedb4a4bd547368b8e8a9076b980d5c83c8.

    Why the previous prototype still failed: it did pad actual tensors in outer Python code, and a real three-rank native probe passed 72 checks. However, model compilation traced that padding branch with unsplit profile metadata. A native A2E observer then caught live tensors with 738 rows against physical control count 802, and 433 against 994. The compiled model was still passing the unpadded payload.

    Final repair: one production file, camp2p.py. Create independent physical stage metadata using the existing strided receiving groups; perform actual padding inside the opaque custom op so it reads current live metadata. Preserve original Attention metadata, query lengths, and E2A reference. The native binding correctly returns the original-length valid prefix. Request-boundary splitting and send → yield → recv remain active.

    Verification Result
    Related connector/token-ID units 24 passed, no skips
    Native A2E/E2A, 3 ranks / 2 cooperative stages / 4 layers 72/72 layout and content checks
    Complete V2-Lite, original 2A1F Legacy DBO counterexample 25/25
    Legacy 2A2F DBO regression 25/25
    Frozen GSM8K first 128, 2A1F DBO enabled 41/128, 32.03%
    Same questions and protocol, DBO disabled 40/128, 31.25%

    Accuracy differs by one question (+0.78 pp). Question, prompt, target hashes and protocols match; both accuracy drivers exit 0. There is no output-identity/strict-match gate or additional score threshold.

    The 2A1F functional log records 248 live split records, including eight unequal records, such as [372, 803] / [305, 433]; 2A2F records 250 live splits. These records exclude warmup/capture/replay. Live split requests use the existing MLA eager fallback while retaining real two-stage DBO.

    The 2A2F model scenario and all saved responses passed. Its shell wrapper subsequently exited 2 because I edited the launcher in place while it was waiting for that command. The completed result was independently audited; the failed wrapper record is retained. I replaced and syntax-checked the launcher before the two accuracy runs. This was an acceptance-driver fault, separate from the repaired model path.

    Hardware validation evidence records pinned dependencies, native/source identities, reproduction and raw evidence archive SHA256. The experiment report is kept in the task workspace rather than committed to the repository. Changed-file pre-commit checks pass. The dedicated Guian task is stopped; local backup and persistent storage are preserved.

    No runner, native binding, vendor-kernel or dependency change is required for this repair. GPU, new NPU MRV2 DBO, #429 and performance benchmarking remain outside this issue's repair. PR #430 is merged into the integration branch. #425 remains open and Draft; no merge into main has been performed. Bug-level implementation and verification are complete; integration into main remains tracked by #425.

    2026-10-09 identity correction: the integration commit is now 563164a, and the published source branch tip is cc508a5. Both retain the exact source trees of 0366406 and 2f506ee, respectively. The Git identity and sign-off now match jiaran-king (Zhou ziheng); all recorded hardware results remain unchanged.

  2. jiaran-king commented on Oct 9, 2026

    @jiaran-king
    CollaboratorAuthor

    Implementation and evidence review completed against current #425 head df0a57c.

    Fix 563164a pads live stage tensors inside the opaque CAMP send operation, with separate physical counts. Original Attention query lengths, the valid return prefix, and real two-stage DBO are preserved. The current production/test files match the hardware-validated source.

    Verified retained results: original 2A1F concurrent counterexample 25/25, Legacy 2A2F regression 25/25, native wire checks 72/72, related units 24 passed. The same-question 128 comparison is DBO on 41/128, off 40/128. The previously disclosed 2A2F shell-wrapper error does not alter its independently verified model results.

    Closing as completed per the requested bug scope. The fix is included in #425's integration branch; final merge into main remains tracked by #425.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions