Repository navigation
feat: support python dspark acl graphs with data parallelism. - #2473
Merged
XuZhang99 merged 1 commit intoOct 11, 2026
Merged
Conversation
Coordinate whole-query padding, graph keys, and capture presence across DP ranks. Materialize safe block-draft inputs for idle peers without writing real cache slots, and cover asymmetric batches and executor selection with unit tests.
PixelFlat
requested review from
DongheJin,
DragonFive,
JimHsiung,
Kang-Meng,
XuZhang99,
liutongxuan,
xiao-yu-chen,
yingxudeng and
zhang-minchao
as code owners
October 11, 2026 15:22
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
XuZhang99
approved these changes
Oct 11, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Enable Python DSpark ACL graphs with data parallelism, building on the graph
replay fixes already merged in #2471. This PR contains only the DP feature.
to a common whole-query request bucket.
asymmetric local batches and page-table capacities.
counts and using invalid cache slots. Clear padding when local batches shrink.
Validation
python setup.py build --device npu --tilelang-jobs 4: passed on the rebased commit.SpecDecodeInputBuilderTestCTest cases: passed.python_tests.*andpython_distributed_tests.*: passed (65 Python modules and three distributedNZ tests), including graph, empty-DP, model-executor, and prepared-graph tests.
--output-on-failure; distributed tests usedautomatically assigned socket ports. This validation does not claim a full
native C++ test-suite run.
Concurrent-request functional validation
Using the same DP=2 graph-enabled configuration, warmup was followed by three
repetitions each at concurrency 2 and 3 (six and nine measured requests,
respectively). All 15 requests completed successfully with the expected 1,868
input tokens, 128 output tokens, and finish reasons; no OOM, timeout, or server
error was observed. These concurrent scenarios validate request completion and
functional behavior, not full numerical output parity or a high-concurrency
throughput target.
Single-request benchmark
GLM-5.2 static W8A8 target with GLM-5.2-DSpark-NPU-0805 draft, Python execution,
16 NPUs, DP=2 / attention TP=8 / EP=1, 7 speculative tokens, ACL graphs and
asynchronous scheduling enabled. Memory utilization is unchanged at 0.88;
temperature is explicitly 0. Each measured request has 1,868 input tokens and
128 output tokens. Results are means of five runs after warmup.
Runtime measurements were obtained on the feature before rebase
(
ee795338); rebase leaves the feature patch unchanged. Acceptance rate is themean of per-request rates; accepted length is total accepted tokens divided by
total draft rounds. This is a functional/performance smoke test, not a
before/after performance comparison.
Summary by CodeRabbit