Skip to content

Repair B200 benchmark correctness, runtime lifecycle, and validation coverage - #21

Merged
cfregly merged 66 commits into
mainfrom
codex/b200-results-hardening-20260906
Sep 7, 2026
Merged

cfregly merged 66 commits into
mainfrom
codex/b200-results-hardening-20260906

Conversation

@cfregly

@cfregly cfregly commented Sep 6, 2026 •

Copy link
Copy Markdown
Owner

This combines the KV/RoPE optimizations from #20 with the direct B200 repair follow-up. Full-output checks exposed CUDA graph ownership, distributed seed/reference ordering, partial DDP accumulation, and live-input verification bugs. The repaired paths preserve the caller workload and retain numerical failures and performance regressions in their reports.

The change also fixes timed allocations, native worker and profiler cleanup, throughput parsing, test seed ownership, and the pinned vLLM/FlashAttention 4 import incompatibility. The vLLM backport is an explicit, hash-checked wheel build/install step. Python and local torchrun entrypoints remain supported without Slurm. README generation preserves the tested commands and their evidence boundaries.

Validation:

  • 486 available targets attempted directly on one or two B200s; skips, informational exits, accuracy-policy gaps, and speed failures are classified separately.
  • The integrated 5,830-case GPU run completed with 5,722 passed, 78 skipped, and 30 failures. After repairs, all 813 tests across the affected files passed on B200. All six isolated Llama tests passed; the standalone MoE entrypoint also passed. The integrated run and its shutdown cleanup remain retained, rather than relabeled as a full-suite pass.
  • Full-model Llama, NanoChat, and transfer optimizations have repeated equivalent-workload measurements and profiler evidence. Neutral and regressing measurements remain visible.
  • 932 benchmark entrypoints passed contract checks with zero errors and warnings; repository-wide Ruff correctness checks and all 12 README checks passed.
  • All 67 profiler regressions passed on CPU and B200. Application-range command builders now request exactly the five validated metrics without adding a section set. A complete baseline report and the repository-built optimized capture passed strict provenance/range/metric inspection. A repeated baseline capture still hit its ten-minute bound; Nsight/NCCL replay remains intermittent and is not claimed stable.
  • Final CI initially completed 5,845 cases with 5,306 passed, 535 skipped, and four worker-import failures. The launcher now uses module execution from the code root with inherited PYTHONPATH removed. All four CPU cases and all eight CPU/CUDA cases on B200 passed, including two-GPU execution. The complete CI retry passed on d77181e: 5,310 CPU tests passed and 535 skipped, with no failures or errors. Static analysis, dashboard checks, and dual-architecture CUDA builds also passed.

The repair checkpoint records each source revision and disposition. Missing model engines, hardware beyond two B200s, approximate-kernel accuracy policies, and incomplete captures remain explicit. These portable observations do not qualify every hardware target or establish that every optimization is faster.

Merged as 8148660, with an identical tree to the CI-tested head. A fresh integrated 5,845-case GPU suite is running from that merged source; its outcome remains pending.

@cfregly cfregly changed the title Enable monolithic CUDA graph capture and preserve memory regressions Repair B200 benchmark correctness, runtime lifecycle, and validation coverage Sep 7, 2026
@cfregly
cfregly changed the base branch from codex/b200-cache-copy-optimization-20260906 to main September 7, 2026 15:51
@cfregly
cfregly marked this pull request as ready for review September 7, 2026 16:20
@cfregly
cfregly merged commit 8148660 into main Sep 7, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant