Purpose
This issue is the execution-level index for active afd-plugin work.
Project direction, support policy, and long-term workstreams remain in [RFC]: afd-plugin project roadmap #155 .
Detailed design and acceptance criteria remain in the linked RFCs and issues.
The linked issue or PR is the source of truth for status; this checklist is a
maintainer-facing snapshot.
Last reviewed: 2026-09-14
Upstream alignment
ModelRunnerV2
Native MoERunner forward refactor
Phase 1: Attention-side native MoERunner reuse
Phase 2: Model-independent FFN MoE computation
Detailed architecture, validation requirements, and acceptance criteria remain
in #225 .
Model support
DeepSeek V4
GLM5.3-flash
Qwen3 MoE
Qwen3.5 / Qwen3.6 MoE
Model performance
Connectors and runtime optimization
Reliability and CI
Merged reliability fixes (FFN graph replay, NPU teardown, and multi-host
initialization): #270 , #279 , #328 .
CI now includes L4 Kubernetes jobs, weekly CUDA E2E, and strengthened
pre-commit checks: #299 , #268 , #271 .
NPU runtime qualification for vLLM 0.28.0 remains tracked in #341 .
Platform and contributor ecosystem
Purpose
This issue is the execution-level index for active afd-plugin work.
maintainer-facing snapshot.
Last reviewed: 2026-09-14
Upstream alignment
ModelRunnerV2
Native MoERunner forward refactor
MoERunnerinjection — [RFC]: Refactor AFD MoE forward around native MoERunner injection #225Phase 1: Attention-side native MoERunner reuse
role-specific runners that share one remote-experts implementation.
selection path.
behavior, and role-aware weight allocation.
parity across their currently supported execution modes.
Phase 2: Model-independent FFN MoE computation
model-specific FFN helper methods.
FusedMoEmodel.MoE exceptions.
Detailed architecture, validation requirements, and acceptance criteria remain
in #225.
Model support
DeepSeek V4
P2pNcclAFDConnectorimplementation — V0.26.0 support for dsv4 (gpu) #191CAMAsyncAFDConnector— [DeepSeek V4][NPU] Support CAMAsyncAFDConnector #227CAMP2pAFDConnector— [DeepSeek V4][NPU] Support CAMP2pAFDConnector #228GLM5.3-flash
CAMAsyncAFDConnector(prioritize this) @yujuancao07CAMP2pAFDConnectorP2pNcclAFDConnectorQwen3 MoE
Qwen3.5 / Qwen3.6 MoE
Qwen3_5MoeForConditionalGenerationfamily — [RFC] Experimental Qwen3.5 MoE AFD support: CUDA correctness baseline #179, Add Qwen3.6 MoE CUDA AFD implementation support #181, test(e2e): add Qwen3.6 MoE CUDA coverage #256Model performance
— [RFC]: DeepSeek V3.2 prefill performance study and token-balanced TP/SP dual batching #170 @ShwStone
Connectors and runtime optimization
Reliability and CI
Merged reliability fixes (FFN graph replay, NPU teardown, and multi-host
initialization): #270, #279, #328.
CI now includes L4 Kubernetes jobs, weekly CUDA E2E, and strengthened
pre-commit checks: #299, #268, #271.
NPU runtime qualification for vLLM 0.28.0 remains tracked in #341.
Platform and contributor ecosystem