Skip to content

Provide recipes and a skill for multi-node deployment on Kubernetes clusters - #421

Open
shlomitk1 wants to merge 15 commits into
vllm-project:mainfrom
ms-llmd:add-multi-node-deployment
Open

shlomitk1 wants to merge 15 commits into
vllm-project:mainfrom
ms-llmd:add-multi-node-deployment

Conversation

@shlomitk1

@shlomitk1 shlomitk1 commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

This PR adds a new Claude Code agent skill, deploy-afd-k8s, plus two new recipe scripts that exercise it. Together they let a user deploy an AFD GPU recipe — including multi-pod placements — onto Kubernetes/OpenShift without manually writing manifests or being asked how to split the workload across pods.

The skill supports arbitrary multi-pod recipe placements (reading topology from the recipe header instead of assuming single-pod), and adds two recipes — one simple 2-pod DBO case and one more complex 3-pod attention-DP-split case — to serve as working examples/validation of that multi-node capability.

Issue

Scope

  • In scope:

    • GPU P2pNcclAFDConnector recipes, prefill_decode_colocation only

    • Multi-node deployment

  • Out of scope:

    • NPU recipes — skill is explicitly GPU-only.

    • prefill_decode_disaggregation recipes

    • E2E correctness testing — explicitly deferred to the separate run-e2e skill

Test Plan

  1. Use the skill with any added recipe, e.g. "Follow .agents/skills/deploy-afd-k8s/SKILL.md to deploy the recipe recipe/gpu/P2pNcclAFDConnector/qwen3_5_122b_a10b_fp8/prefill_decode_colocation/multipod_2a_2a_2f_graph.sh and expose the endpoint"

  2. Send some prompt to the provided endpoint.

Test Result

Both new multi-pod recipes were deployed end-to-end on a live OpenShift
cluster (image afd-ci:latest built using docker/Dockerfile.ci), one at a time, with the
DeepSeek deployment torn down before starting the Qwen one to free GPUs.

1. DeepSeek-V2-Lite — multipod_2a_2f_graph_dbo_dp1tp2.sh (2 pods)

Pod → node placement:

Pod Node Role
vllm-pod-attention-0 ...44s2 afd-attn-node-role=head
vllm-pod-ffn-0 ...44s2 afd-ffn-node-role=head

Readiness logs:

[attn] --- waiting on attn.log (marker: Application startup complete) ---
[attn] === pod vllm-pod-attention-0 READY ===
[ffn]  --- waiting on ffn.log (marker: AFD FFN EngineCore started) ---
[ffn]  === pod vllm-pod-ffn-0 READY ===

kubectl get pods showed 1/1 Running for both — confirming the
readinessProbe passes.

Inference: POST /v1/chat/completions,
{"messages":[{"role":"user","content":"What is the capital of France?"}], "max_tokens":64, "temperature":0}
→ "Paris.\n\nUser: What is the capital of the United States?..."
(finish_reason: length, 64 completion tokens).

Torn down afterward (kubectl delete pod vllm-pod-attention-0 vllm-pod-ffn-0,
kubectl delete svc vllm-ffn-p2p-service) to free GPUs.

2. Qwen3.5-122B-A10B-FP8 — multipod_2a_2a_2f_graph.sh (3 pods)

Pod → node placement (real cross-node spread, no anti-affinity requested):

Pod Node Role
vllm-pod-attention-0 ...44s2 afd-attn-node-role=head
vllm-pod-attention-1 ...43s0 afd-attn-node-role=worker
vllm-pod-ffn-0 ...44s2 afd-ffn-node-role=head

Readiness logs (per pod):

[attn-0] --- waiting on attn.log (marker: Application startup complete) ---
[attn-0] --- waiting on attention server /health (127.0.0.1:18305) ---
[attn-0] attention server: ready
[attn-0] === pod vllm-pod-attention-0 READY ===

[attn-1] --- waiting on attn.log (marker: init engine (profile, create kv cache, warmup model) took) ---
[attn-1] === pod vllm-pod-attention-1 READY ===

[ffn-0]  --- waiting on ffn.log (marker: AFD FFN EngineCore started) ---
[ffn-0]  === pod vllm-pod-ffn-0 READY ===

All three show 1/1 Running.

Inference: same request as above →
"Thinking Process: ... The capital of France is Paris. ..."
(finish_reason: length, system_fingerprint: vllm-0.26.0-dp4-ep-...,
64 completion tokens).

Docs Impact

recipe/gpu/P2pNcclAFDConnector/README.md


Essential PR Checklist
  • Purpose is clear and linked to public context when possible.
  • Scope is bounded.
  • Compatibility with vLLM v0.26.0 is considered.
  • No changes are made to the vLLM source checkout.
  • Plugin-owned classes or explicit dotted class paths are preferred over monkey patches.
  • Any compat shim or monkey patch is isolated, idempotent, version-guarded, documented, and tested.
  • Imports remain CPU-safe; CUDA-heavy work is delayed or GPU-gated.
  • Validation evidence is included, including skipped GPU tests when applicable.
  • Documentation impact is stated.

Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
@hsliuustc0106 hsliuustc0106 added enhancement New feature or request documentation Improvements or additions to documentation labels Oct 6, 2026
fi

echo "=== pod ${TEMPLATE_POD} READY ==="
sleep infinity

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Keep Kubernetes readiness tied to the serving processes

After the initial startup checks, this shell sleeps indefinitely without monitoring the recipe or its Attention/FFN processes. If serving exits later, the container remains running; without a readiness probe, the client Service can continue routing requests to the failed instance. The fail() path also sleeps indefinitely. Please monitor the serving processes and exit on failure, and add a readiness probe that reflects serving health.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • Replaced the terminal sleep infinity with a monitoring loop that polls the recipe process every 10s (and the /health endpoint for attention-head pods), calling fail() the moment serving dies.
  • Changed fail() to remove the /work/ready marker and exit 1 instead of hanging forever — a crashed container now actually terminates (pod goes Failed under restartPolicy: Never), while still tailing logs first for diagnosis.
  • Added a readinessProbe (exec) checking the /work/ready marker, the recorded recipe.pid, and (for head pods) /health — so vllm-service, which only routes to afd-attn-node-role: head pods, stops sending traffic the moment serving becomes unhealthy, independent of the container-exit path.

Upstream verifies cross-node `P2pNcclAFDConnector` placement specifically for
DeepSeek-V2-Lite `2A2F` with the FFN ranks on separate hosts, over TCP (ENA)
and over EFA; every other topology -- including the qwen3.5 3-pod plan above
-- is still **unverified**

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Add deployment evidence for the new multi-pod recipes

This document explicitly marks the Qwen three-pod topology as unverified, and templates/pod.yaml also says the headless-worker readiness marker has not been observed on a live run. These are behaviors the new deployment workflow relies on, while the PR test result provides no concrete execution evidence. Please include the commands, image/version, actual Pod-to-node placement, startup/readiness logs for each role, and a successful inference request for the new recipes. At minimum, verify the DeepSeek two-pod recipe across nodes; if the Qwen recipe cannot be exercised yet, clearly label it as an experimental example rather than a validated deployment.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Edited the document to remove obsolete "unverified" comments. Both recipes have been successfully deployed. I've attached the evidence to the PR description under "Test result".

Signed-off-by: Shlomit Koyfman <shlomitk@il.ibm.com>
@jiangkuaixue123

Copy link
Copy Markdown
Collaborator

The main branch has now been updated to v0.30.0. Please sync this PR with the latest main and resolve any merge conflicts before merging.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[RFC]: Provide recipes and a skill/documentation for multi-node deployment on Kubernetes clusters

3 participants