Repository navigation
[Bug]: NPU MRV1 2A1F Legacy DBO corrupts outputs with unequal live stage lengths #428
Description
Activity
The repair from PR #430 has been merged into #425's integration branch as
563164a. The merged tree is identical to the reviewed and hardware-validated #430 tip2f506eedb4a4bd547368b8e8a9076b980d5c83c8.Why the previous prototype still failed: it did pad actual tensors in outer Python code, and a real three-rank native probe passed 72 checks. However, model compilation traced that padding branch with unsplit profile metadata. A native A2E observer then caught live tensors with 738 rows against physical control count 802, and 433 against 994. The compiled model was still passing the unpadded payload.
Final repair: one production file,
camp2p.py. Create independent physical stage metadata using the existing strided receiving groups; perform actual padding inside the opaque custom op so it reads current live metadata. Preserve original Attention metadata, query lengths, and E2A reference. The native binding correctly returns the original-length valid prefix. Request-boundary splitting andsend → yield → recvremain active.Verification Result Related connector/token-ID units 24 passed, no skips Native A2E/E2A, 3 ranks / 2 cooperative stages / 4 layers 72/72 layout and content checks Complete V2-Lite, original 2A1F Legacy DBO counterexample 25/25 Legacy 2A2F DBO regression 25/25 Frozen GSM8K first 128, 2A1F DBO enabled 41/128, 32.03% Same questions and protocol, DBO disabled 40/128, 31.25% Accuracy differs by one question (+0.78 pp). Question, prompt, target hashes and protocols match; both accuracy drivers exit 0. There is no output-identity/strict-match gate or additional score threshold.
The 2A1F functional log records 248 live split records, including eight unequal records, such as
[372, 803]/[305, 433]; 2A2F records 250 live splits. These records exclude warmup/capture/replay. Live split requests use the existing MLA eager fallback while retaining real two-stage DBO.The 2A2F model scenario and all saved responses passed. Its shell wrapper subsequently exited 2 because I edited the launcher in place while it was waiting for that command. The completed result was independently audited; the failed wrapper record is retained. I replaced and syntax-checked the launcher before the two accuracy runs. This was an acceptance-driver fault, separate from the repaired model path.
Hardware validation evidence records pinned dependencies, native/source identities, reproduction and raw evidence archive SHA256. The experiment report is kept in the task workspace rather than committed to the repository. Changed-file pre-commit checks pass. The dedicated Guian task is stopped; local backup and persistent storage are preserved.
No runner, native binding, vendor-kernel or dependency change is required for this repair. GPU, new NPU MRV2 DBO, #429 and performance benchmarking remain outside this issue's repair. PR #430 is merged into the integration branch. #425 remains open and Draft; no merge into main has been performed. Bug-level implementation and verification are complete; integration into main remains tracked by #425.
2026-10-09 identity correction: the integration commit is now
563164a, and the published source branch tip iscc508a5. Both retain the exact source trees of0366406and2f506ee, respectively. The Git identity and sign-off now matchjiaran-king(Zhou ziheng); all recorded hardware results remain unchanged.- added 2 commits that reference this issue
on Oct 8, 2026 Implementation and evidence review completed against current #425 head
df0a57c.Fix
563164apads live stage tensors inside the opaque CAMP send operation, with separate physical counts. Original Attention query lengths, the valid return prefix, and real two-stage DBO are preserved. The current production/test files match the hardware-validated source.Verified retained results: original 2A1F concurrent counterexample 25/25, Legacy 2A2F regression 25/25, native wire checks 72/72, related units 24 passed. The same-question 128 comparison is DBO on 41/128, off 40/128. The previously disclosed 2A2F shell-wrapper error does not alter its independently verified model results.
Closing as completed per the requested bug scope. The fix is included in #425's integration branch; final merge into main remains tracked by #425.
- added a commit that references this issue
on Oct 9, 2026
Summary
During #425 NPU acceptance, synchronous MRV1 Legacy DBO on 2A1F returned HTTP
200 but generated corrupt content when real concurrent requests produced
unequal per-stage Attention token counts. The historical-recipe repeat passed
only 12/25 content checks: the serial request passed, but 13 of the 24
concurrent responses were garbled. DBO-disabled 2A1F eager/graph and Legacy
DBO 2A2F passed the same 25-response probe.
This tracks an existing Legacy NPU path. It is separate from new NPU MRV2 DBO
support and does not require taking over #422. Repair work resumed after the bounded Goal was approved. A verified fix is
now published in PR #430,
merged into #425's integration branch as
563164a.The implementation and recorded device evidence have been reviewed against
current #425 head
df0a57c. The reported bug is fixed and bug-level validationis complete. Main integration remains tracked in #425, which is still open
and Draft; the fix has not yet merged into main.
Current environment
3ee2e812656a0d88baf6487b02fcf6f15a88ea2e; the affected Legacy runner isretained in follow-up
605ba33ac27a4fc667d5d4ef93732d9c5a4f626c.ced6857afa0ea7b2e3f0846a62e1394e90f15607.8d4409d6256d8a6729140ddcc0d1889e3f96cdd6.CANN 9.1.0; torch 2.10.0+cpu; torch-npu 2.10.0.post4; driver 25.5.1.1.
are unchanged.
Reproduction
Use the existing NPU
afd-graph-dbo-2a1fscenario with Attention DP2/TP1,one FFN rank, block size 128, max model length 4096, max batched tokens 4096,
max sequences 8, prefix caching and chunked prefill disabled, and DBO
decode/prefill thresholds 2/8.
The acceptance workload is one serial completion followed by 24 completions
at concurrency 12. Four prompt lengths are 305, 369, 433, 497 tokens.
The following driver reuses the repository launcher and its existing accounting
smoke prompt; save it as
reproduce_legacy_2a1f.pyin the repository root:The device evidence was collected using this workload through an acceptance
wrapper around the same launcher. The standalone packaging above has not been
rerun after repair work was paused. Check live stage records, not just DBO
configuration or capture records.
Actual behavior and confirmed evidence
!,zero-width characters and unrelated multilingual tokens. All four prompt
lengths have failures; this is not a GSM8K extraction or strict-match issue.
[802, 433] and stage 1 [866, 674]. Decode also produced [3, 2]
and [4, 3].
is_graph_capturing=False,is_warmup=False,is_graph_replaying=False.The existing MLA split fallback is already active: corruption occurs during
eager split execution.
receives a total count of 1235 and divides it into 617/617, while the A
payloads contain 802/433 rows. Both row coverage and payload offsets disagree.
The equal-chunk communication contract is demonstrably violated. Explaining
the observed corruption by this mismatch is a source-based diagnosis, not a
claim that a unique root cause or successful repair has been established.
An isolated connector stage-padding prototype passed 20 related unit cases
but still passed only 10/25 device content checks. It is uncommitted and
unpublished. A subsequent observational probe was interrupted at the
requester's instruction and has no completion verdict. These results must not
be presented as a fix or as evidence that metadata padding alone is sufficient.
History and source pointers
v0.28 qualification. Commit
1b4ca7benabled whole-batch CAMP fan-in padding. The whole-batch protection is still
present in the current runner.
documented that those runs split only during warmup/capture, not live
requests. They do not establish earlier correctness of this live-split case.
18ef89d4introduced request-boundary splitting into synchronous eager DBO. It can
turn an equal padded batch into unequal physical stages.
CAMP payload handling,
A2E kernel,
E2A kernel.
Expected behavior and bounded follow-up
Existing Legacy 2A1F DBO should return correct contents for unequal real
workloads. A repair must reconcile per-stage physical send/receive sizes with
the kernel protocol while preserving MLA's true query lengths. Changing only
control counts, padding only the whole batch, or reapplying the existing graph
fallback does not establish a solution.
Acceptance should rerun this concurrent counterexample, demonstrate real
two-stage execution and correct physical sizes, and retain a Legacy 2A2F
regression check. Use a bounded same-question accuracy comparison if the
numerical path changes; no full 1319-question rerun or strict-match gate is
required. If a fallback is used instead, disclose the lost DBO overlap rather
than labelling it preserved DBO support.
Before submitting
Commit identity was corrected on 2026-10-09 to
jiaran-king(Zhou ziheng). Integration commit563164areplaces0366406with an identical source tree; hardware results and the archived source identities are unchanged.