Skip to content

feat(infra): training image + GPU Job runbook for ML squad - #30

Merged
Enskc05 merged 1 commit into
mainfrom
devops/feature
Aug 7, 2026
Merged

Enskc05 merged 1 commit into
mainfrom
devops/feature

Conversation

@Enskc05

@Enskc05 Enskc05 commented Aug 7, 2026

Copy link
Copy Markdown
Collaborator

feat(infra): training image + GPU Job runbook

The ML squad could not run training on the server: infra/docker/README.md
listed training.Dockerfile, but the file did not exist. Without an image,
the training code cannot run as a Kubernetes Job.

Changes

  • infra/docker/training.Dockerfile — built on the stock PyTorch CUDA
    image rather than pip-installing torch (avoids a multi-GB download and the
    risk of a CUDA/driver version mismatch). Runs as a non-root user; only
    services/ml is copied into the image.
  • infra/docker/training.Dockerfile.dockerignore — limits the build
    context to requirements/ and services/ml/, keeping the Go API and the
    Next.js frontend out of the build.
  • requirements/ml.txt — added hydra-core, omegaconf, boto3, and
    python-dotenv. services/ml/training/train.py and
    services/ml/minio_loader.py import these, but they were missing from the
    list: the image would build fine and the first run would fail with
    ImportError.
  • docs/runbooks/training-job.example.yaml — GPU Job template. Includes
    an emptyDir mounted at /dev/shm; the container default is 64 MB, which
    PyTorch DataLoader workers exhaust, crashing with Bus error.
  • docs/runbooks/ML-TRAINING.md — one-time DevOps setup (image build and
    push, MinIO SealedSecret) plus the ML squad's day-to-day run loop and a
    troubleshooting table.

Notes

Python version. The image ships Python 3.11 (stock PyTorch CUDA image)
while the repo root requires 3.13. This divergence is deliberate: th
floor exists for ehtim (data squad) and the training code does not use it.
The rationale is documented in the Dockerfile header.

Image tags. Tags are date-based (2026-08-07); latest is deliberately
avoided because Kubernetes will not re-pull an unchanged tag, so an image
update would silently fail to apply.

Question for the ML squad

services/ml/conf/data/default.yaml sets bucket_name: datasets and
minio_prefix: datasets/training-512/v1, while the layout in docs/ is a datasetsbucket with atraining-512/v1/prefix. If both are correct as written, boto3 will look fordatasets/datasets/training-512/v1nothing. Worth verifying withmc ls dh/datasets/` before the first run.

Testing

Not yet built on the server — the training code lives on `ml/feature
is not merged. Build recipe is in the runbook; the Dockerfile's fina
step imports every required package and asserts that torch was built
CUDA, so a missing dependency fails the build rather than the third
training run.
main...devops/featur

PR açıklaması (kopyala)

▎ feat(infra): training image + GPU Job runbook
▎
▎ ML ekibinin sunucuda eğitim koşabilmesi için eksik olan parçalar. ining.Dockerfile'ı listeliyordu ama dosya yoktu — kod bir konteyneregirmeden Job olarak koşturulamıyordu.
▎
▎ - training.Dockerfile — stok PyTorch CUDA imajı tabanlı (torch pip ile kurulmadı: GB'larca indirme + CUDA sürüm eşleşmesi riski). Non-root, sadece services/ml kopyalanıyor.
▎ - requirements/ml.txt — hydra-core, omegaconf, boto3, python-dotenv eklendi. train.py ve minio_loader.py bunları import ediyordu ama listede yoktular; imaj build
▎ olur, ilk koşu ImportError ile düşerdi.
▎ - training-job.example.yaml — GPU Job şablonu. /dev/shm için emptyDir: varsayılan 64 MB, PyTorch DataLoader worker'larını Bus error ile düşürüyor.
▎ - ML-TRAINING.md — DevOps'un bir kerelik kurulumu + ML squad'ın ko.
▎
▎ Not: İmajın Python'u 3.11 (stok PyTorch imajı), repo kökü 3.13 istiyor. Fark bilinçli — 3.13 alt sınırı ehtim için, eğitim kodu onu kullanmıyor. Gerekçe Dockerfile başlığında.
▎
▎ ML squad'a soru: conf/data/default.yaml bucket_name: datasets ve minio_prefix: datasets/training-512/v1 diyor. docs/DATA.md düzeni datasets bucket'ı +
▎ training-512/v1/ prefix'i. İkisi birlikteyse boto3 datasets/datasence doğrulanmalı.

Merge sonrası sunucuda build için origin/devops/feature yerine origin/main'den alabilirsin:

cd ~/deephorizon && git fetch origin
git checkout -B build/training origin/ml/feature
git checkout origin/main -- infra/docker/training.Dockerfile infra/docker/training.Dockerfile.dockerignore requirements/ml.txt

TAG=$(date +%F)
docker build -f infra/docker/training.Dockerfile -t localhost:32000/deephorizon-training:$TAG .

Build çıktısında takılırsan logu at, devops/fix ile düzeltiriz.

✻ Sautéed for 25s

❯ PR açıklamasını ingilizce ve .md formatta yaz

feat(infra): training image + GPU Job runbook

The ML squad could not run training on the server: infra/docker/REA listed training.Dockerfile, but the file did not exist. Without an image, the training code cannot run as a Kubernetes Job`.

Changes

  • infra/docker/training.Dockerfile — built on the stock PyTorch CUDA
    image rather than pip-installing torch (avoids a multi-GB download and the
    risk of a CUDA/driver version mismatch). Runs as a non-root user;
    services/ml is copied into the image.
  • infra/docker/training.Dockerfile.dockerignore — limits the b
    context to requirements/ and services/ml/, keeping the Go API and the
    Next.js frontend out of the build.
  • requirements/ml.txt — added hydra-core, omegaconf, boto3, and
    python-dotenv. services/ml/training/train.py and
    services/ml/minio_loader.py import these, but they were missing from the
    list: the image would build fine and the first run would fail with
    ImportError.
  • docs/runbooks/training-job.example.yaml — GPU Job template
    an emptyDir mounted at /dev/shm; the container default is 64 MB, which
    PyTorch DataLoader workers exhaust, crashing with Bus error.
  • docs/runbooks/ML-TRAINING.md — one-time DevOps setup (image build and
    push, MinIO SealedSecret) plus the ML squad's day-to-day run loop
    troubleshooting table.

Notes

Python version. The image ships Python 3.11 (stock PyTorch CUDA image)
while the repo root requires 3.13. This divergence is deliberate: the 3.13
floor exists for ehtim (data squad) and the training code does not use it.
The rationale is documented in the Dockerfile header.

Image tags. Tags are date-based (2026-08-07); latest is deliberately
avoided because Kubernetes will not re-pull an unchanged tag, so an
update would silently fail to apply.

Question for the ML squad

services/ml/conf/data/default.yaml sets bucket_name: datasets and
minio_prefix: datasets/training-512/v1, while the layout in docs/ is a datasetsbucket with atraining-512/v1/prefix. If both are correct as written, boto3 will look fordatasets/datasets/training-512/v1and find nothing. Worth verifying withmc ls dh/datasets/` before the first

Testing

Not yet built on the server — the training code lives on ml/feature, which
is not merged. Build recipe is in the runbook; the Dockerfile's final RUN
step imports every required package and asserts that torch was built
CUDA, so a missing dependency fails the build rather than the third
training run.

Egitim kodu (services/ml/training) bir konteynere girmeden K8s Job olarak
kosturulamiyordu; infra/docker/README'de listelenen training.Dockerfile
mevcut degildi.

- infra/docker/training.Dockerfile: stok pytorch CUDA imaji (torch+cuDNN
  hazir) uzerine requirements/ml.txt. Yalnizca services/ml kopyalanir,
  non-root calisir, Hydra ciktilari icin yazilabilir /workspace.
- requirements/ml.txt: hydra-core, omegaconf, boto3, python-dotenv eklendi.
  Egitim giris noktasi bunlari import ediyordu ama listede yoktular; imaj
  build oluyor, ilk kosu ImportError ile duserdi.
- docs/runbooks/training-job.example.yaml: GPU Job sablonu. /dev/shm icin
  emptyDir — varsayilan 64 MB PyTorch DataLoader worker'larini "Bus error"
  ile dusuruyor.
- docs/runbooks/ML-TRAINING.md: DevOps'un bir kerelik kurulumu (imaj build,
  MinIO SealedSecret) + ML squad'in kosu dongusu + sorun giderme tablosu.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@Enskc05
Enskc05 merged commit 2e33988 into main Aug 7, 2026
0 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant