Make benchmark execution declarative instead of script-by-script, while preserving the repo's existing artifact and methodology contracts.
The service contract is a BenchmarkRun custom resource:
- workload definition
- benchmark layer selection
- comparison variable
- scheduler path
- measurement policy
- provenance policy
- publication-grade versus realism-grade mode
The CRD lives at cluster/configs/benchmarkrun-crd.yaml. The local example lives at templates/benchmark_run.yaml.
Use templates/performance_intake.yaml plus templates/benchmark_workload_spec.yaml to freeze the human-reviewed workload contract before generating or submitting a BenchmarkRun.
Use docs/performance_warehouse.md plus templates/performance_warehouse_contract.yaml to keep raw evidence, warehouse lineage, and retention policy consistent across operator implementations.
- A user, CI workflow, or release process submits a
BenchmarkRun. - The operator validates that the workload is frozen, the workload spec is pinned, and only one comparison variable is changing.
- The operator resolves the requested layer set:
- micro
- component
- end-to-end
- The operator chooses execution primitives:
- SLINKY for topology-aware placement and isolation policy
- Kueue for queueing, admission, cohorting, and canary/nightly/pre-release scheduling
- The operator launches the matching repo entrypoints:
bench run-tier1cluster common-eval- targeted cluster scripts for NCCL, fio, vLLM sweeps, and startup checks
- Raw artifacts land in immutable object storage plus the run directory contract already used by the repo.
- Results are copied into:
- a hot time-series system for operational alerting and debugging
- a long-term analytical warehouse for regression analysis across releases, hardware generations, and data centers
- The operator updates
status.conditions, links artifact locations, and records audit metadata.
The operator should treat the warehouse as part of the product surface, not an afterthought.
Required split:
- raw logs, traces, profiler outputs, manifests, and exported metrics in immutable low-cost storage
- reduced-cardinality operational metrics in a hot time-series system
- curated fact and dimension tables in a columnar analytical store
Required curated dimensions:
- software versions
- hardware topology
- cluster and region metadata
- workload parameters
- artifact lineage
Required curated facts:
- benchmark run facts
- serving outcomes
- training outcomes when applicable
- telemetry slices for comparison and regression detectors
BenchmarkRun.spec.observability should give the operator enough structure to join service, job, and hardware telemetry without ad hoc heuristics.
Required signals include:
- service metrics and request traces
dcgm-exporternvlink-exporternode-pci-exporterping-exporternode-problem-detectorhpc-verification- Kubernetes, Kueue, and SLINKY scheduling signals
Required stable join keys include:
run_idbenchmark_case_idscheduler_run_idjob_uidpod_uidnode_namegpu_uuidrank_idfor distributed pathsrequest_idandtrace_idfor serving paths
The service should make these questions cheap to answer from the warehouse and raw evidence:
- inference p99 doubled but average barely moved
- training throughput dropped while GPU utilization stayed low
- distributed runs show periodic spikes and one node looks slow
- jobs spend too long pending and cluster utilization regresses
Those scenarios are modeled directly in spec.observability.scenarioPlaybooks so the operator and downstream warehouse know which pivots must remain intact.
| Field | Why it exists |
|---|---|
spec.layers |
Expresses the micro/component/end-to-end stack explicitly. |
spec.workload |
Freezes model, sequence-length mix, precision, batching policy, and concurrency model. |
spec.comparison.variableUnderTest |
Forces one-variable-at-a-time comparisons. |
spec.metrics |
Separates training and inference success criteria. |
spec.trials |
Declares replicates, confidence level, and outlier policy. |
spec.bottleneckAnalysis |
Binds taxonomy to instrumentation instead of guessing from GPU utilization alone. |
spec.distributed |
Requires rank-level visibility and node diagnosis for distributed runs. |
spec.provenance |
Makes image digests, topology, audit trails, and signed attestations first-class policy. |
spec.executionPolicy |
Separates publication-grade isolation from realism-grade customer-experience testing. |
spec.observability |
Declares telemetry joins, exporter coverage, and required debug playbooks. |
spec.sinks |
Splits raw evidence, hot operational metrics, and long-term regression analytics. |
| Concern | SLINKY role | Kueue role |
|---|---|---|
| Placement | Enforce topology-aware placement, GPU/NIC affinity, exclusive placement for publication-grade runs | Admit the run into the correct queue and cohort |
| Isolation | Choose dedicated or topology-exclusive placement classes | Coordinate admission limits and multi-tenant fairness |
| Cadence | N/A | schedule canary, nightly, and pre-release runs |
| Backpressure | N/A | queue depth, retries, and scheduling policy |
The operator should translate the BenchmarkRun execution mode into scheduler semantics rather than forcing workload owners to understand queueing internals.
Use the existing repo unit of record:
cluster/runs/<run_id>/
manifest.json
structured/
raw/
figures/
reports/
For single-node or non-cluster paths, keep using the existing benchmark history package under artifacts/history/tier1/<run_id>/.
Publication-grade runs should resolve to:
- dedicated nodes
- fixed topology
- stable background load
- topology-exclusive scheduling when appropriate
- image digests, not mutable tags
- signed provenance policy enabled
If those guarantees are unavailable, the operator should mark the run as realism-grade or reject it. It should not silently continue and let the user believe the result is publication-safe.
Realism-grade runs should intentionally model customer conditions:
- multi-tenant nodes when requested
- realistic background load
- cluster context capture preserved
- outliers explained with scheduler, topology, and node telemetry
This mode is not weaker. It answers a different question.
Recommended status additions for a future operator implementation:
status.phasestatus.conditionsstatus.selectedSchedulerPathstatus.runIdstatus.runDirectorystatus.hotMetricsRefstatus.warehouseRefstatus.outlierSummarystatus.stragglerNodesstatus.provenanceRef
Recommended automation policies:
- canary: smallest useful micro or component run after merge
- nightly: tier-1 plus selected cluster components
- pre-release: publication-grade end-to-end suite with full provenance requirements
These schedules should create BenchmarkRun resources or equivalent generated specs, not opaque shell scripts with hidden defaults.