Skip to content

feat: BYO clients ship their own container logs; drop the CLI's detached shipper - #479

Merged
7174Andy merged 2 commits into
mainfrom
andrew/feat-byo-in-container-log-shipper
Aug 26, 2026
Merged

7174Andy merged 2 commits into
mainfrom
andrew/feat-byo-in-container-log-shipper

Conversation

@7174Andy

Copy link
Copy Markdown
Collaborator

Summary

  • BYO clients now ship their own container logs to the allocator from inside the container: start.sh routes fd 5 (every service's tagged output) through a new ship_logs worker in the client package when SHIP_LOGS=1, which lablink client register writes for every BYO shape.
  • Closes the gap where hand-off clients (register --no-run-locally, e.g. Run:AI workloads) had no log shipper at all — the container is its own PID 1 with no docker daemon to tail, so their allocator logs pages stayed permanently empty.
  • Deletes the CLI's detached log_shipper.py (441 lines + 673 test lines), its PID/state-file machinery, Docker.follow_logs, and the psutil dependency — net −1,284 lines.

Changes Made

Client package

  • New lablink_client_service/ship_logs.py + ship_logs console script: reads fd 5 from stdin, writes each line to container stdout first (so docker logs output is byte-identical), queues a UTC-stamped copy in a bounded drop-oldest deque, and ships batches (50 lines / 15 s) to POST /api/vm-logs/<VM_NAME> from a separate thread. SIGTERM/EOF do a single-attempt final flush so a graceful docker stop ships the tail; missing env degrades to pure passthrough.
  • start.sh: when SHIP_LOGS=1, fd 5 feeds a supervisor loop that respawns a crashed worker (the kernel pipe buffer absorbs the gap) and exec cats after 3 failures — worst case is exactly today's plumbing, never a frozen container.

CLI package

  • register writes SHIP_LOGS=1 into the env file unconditionally (run-locally via --env-file, hand-off via the paste block); the shipper spawn/revive/stop machinery is deleted and _resume shrinks to container-only. CONTAINER_NAME moves to register.py.
  • client doctor's shipper check becomes docker exec lablink-client pgrep -f ship_logs (pgrep verified present in the published image).
  • docs/reference/cli.md's "in-container shipper" wording is now literally accurate; docs/cli/byo-clients.md documents the hand-off paste requirement.

Design Decisions

  • Read the source stream, no file intermediary. An earlier tee-to-file design was rejected because it rebuilt the architecture of the fix: resolve VM log shipping data loss from Docker json-file driver stall #304 json-file tail stall (writer freezes silently, shipper tails a dead file forever) and made tee a single point of failure for all container logging.
  • Passthrough-first, fail-open. The worker sits in the logging path, so the read loop never blocks on the network (bounded queue, shipping thread) and every failure mode degrades to plain passthrough — logs visible in docker logs / runai workspace logs, absent from the allocator, loudly visible either way.
  • One shipper for all BYO shapes instead of a hand-off-only patch: the CLI's detached shipper was the flakiest component (a background process on operator laptops that died silently in live testing) and docker logs is the same stream the container sees internally.
  • 4xx drops the batch instead of killing the worker (unlike the old CLI shipper): nothing respawns this worker for shipping-only failures, so a transient 401 during allocator DB warmup must not permanently kill the only log channel.
  • Log group is container-docker — the allocator's existing -docker suffix routing lands it in the docker_logs column with zero allocator changes.

Release ordering

The client image must be published with the ship_logs entry point before a new CLI release reaches operators — a new CLI plus an old image means BYO clients log locally but nothing reaches the allocator (old images simply ignore SHIP_LOGS). AWS clients are untouched (user_data's log_shipper.sh remains the host-side shipper there, including cloud-init logs).

Testing

  • 11 new client tests (test_ship_logs.py): passthrough byte-fidelity and ordering, timestamping, drop-oldest overflow, batch drain, retry/drop semantics, single-retry final flush, interval + stop flushes.
  • CLI tests reworked: SHIP_LOGS=1 asserted in env file and hand-off printout, doctor's new exec-based check, resume path without shipper revival; stale shipper tests deleted.
  • Full suites: client 178 passed; CLI 777 passed (12 pre-existing local-only ANSI rendering failures, unrelated).
  • Live shell simulation of the supervisor: stub worker crashing every 2 lines → all 12 lines passed through exactly once across 3 respawns, then the cat fallback engaged.
  • Real ship_logs binary: clean passthrough with missing env and with an unreachable allocator.

7174Andy and others added 2 commits August 26, 2026 15:46
…hed shipper

BYO log shipping was host-side only: run-locally boxes relied on a
detached CLI process that could die silently on operator laptops, and
hand-off clients (register --no-run-locally, e.g. Run:AI workloads)
had no shipper at all -- the container is its own PID 1 with no docker
daemon to tail, so their allocator logs pages stayed empty.

Now the container ships its own stream. start.sh routes fd 5 (every
service's tagged output) through a new ship_logs worker in the client
package when SHIP_LOGS=1, which register writes for every BYO shape:

* Passthrough first: each line is written to container stdout before
  anything else touches it, so `docker logs` output is byte-identical;
  shipping happens on a separate thread fed by a bounded drop-oldest
  queue and can never block the read loop (the lablink#304 tail-stall
  lesson: no file intermediary, read the source stream).
* Fail open: a supervisor loop respawns a crashed worker and execs
  plain `cat` after 3 failures -- worst case is exactly the old
  plumbing, never a frozen container. Missing env degrades to pure
  passthrough.

The CLI's log_shipper.py (441 lines + 673 test lines), its PID/state
files, register's spawn/revive machinery, Docker.follow_logs, and the
psutil dependency are all deleted. doctor's shipper check becomes a
pgrep inside the container.

Release ordering: the client image must ship with the ship_logs entry
point before a new CLI reaches operators, or BYO clients log locally
but nothing reaches the allocator. Old images ignore SHIP_LOGS.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AzEx3SZp9EAmTVKe6JEWRa
CI's client coverage gate failed at 89% (fail-under=90): every worker
primitive was tested but main() itself — env wiring, the shipper
thread, signal-handler registration, the EOF final flush, and the
missing-env passthrough degrade — was 34 uncovered statements (module
at 66%). Two end-to-end tests through main() bring the module to 96%.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AzEx3SZp9EAmTVKe6JEWRa
@7174Andy
7174Andy merged commit 652a834 into main Aug 26, 2026
11 checks passed
@7174Andy
7174Andy deleted the andrew/feat-byo-in-container-log-shipper branch August 26, 2026 23:51
7174Andy added a commit that referenced this pull request Sep 8, 2026
Release prep. publish-pip.yml's version guardrail rejects a tag whose
version does not match pyproject.toml, so the bumps land on main before
the release tags are cut.

Allocator and client stay in lockstep at 0.4.0 as they have since 0.1.0;
the CLI is versioned independently and goes to 0.3.0.

The CLI's allocator pin is raised to >=0.4.0 this time: the CLI
re-exports MachineConfig, whose ami_id default became empty (= resolve
the per-region Deep Learning Base AMI, #489) in allocator 0.4.0. An
older allocator would silently reintroduce the stale hardcoded
us-west-2 AMI default that doctor's #490 fallback logic assumes gone.

CHANGELOG (CLI): new 0.3.0 section (#484, #485, #490, #491, #498), and
a backfilled 0.2.0 section — #481 tagged 0.2.0 without adding one
(#467, #472, #474, #479).

Also: README Docker <version> example moved to 0.4.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
7174Andy added a commit that referenced this pull request Sep 8, 2026
Release prep. publish-pip.yml's version guardrail rejects a tag whose
version does not match pyproject.toml, so the bumps land on main before
the release tags are cut.

Allocator and client stay in lockstep at 0.4.0 as they have since 0.1.0;
the CLI is versioned independently and goes to 0.3.0.

The CLI's allocator pin is raised to >=0.4.0 this time: the CLI
re-exports MachineConfig, whose ami_id default became empty (= resolve
the per-region Deep Learning Base AMI, #489) in allocator 0.4.0. An
older allocator would silently reintroduce the stale hardcoded
us-west-2 AMI default that doctor's #490 fallback logic assumes gone.

CHANGELOG (CLI): new 0.3.0 section (#484, #485, #490, #491, #498), and
a backfilled 0.2.0 section — #481 tagged 0.2.0 without adding one
(#467, #472, #474, #479).

Also: README Docker <version> example moved to 0.4.0.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant