Skip to content

design-pipeline has no per-child spend or wall-clock bound, and the burn watcher prescribes a remedy for a cause it cannot distinguish #2396

Description

@eumemic

Observation

lane-child-burn-watch fired on run wfr_01M1WEDVR1E7RRZBAX5NZ7PC3A (workflow design-pipeline v5, trigger design-sweep-15m) reporting a suspended run with a live child.

Measured directly against the usage ledger, not run state:

time child events child cost
23:41Z 3,780 $35.04
23:56Z 5,885 $56.42

That is roughly $85/hour and ~2,100 events per 15 minutes, for a grounding pass whose declared output schema is five small arrays (facts, seam_files, substrate_core, half_built_prior_art, thread_directives).

The child is not wedged. Across ~30 sampled tool calls, 22 argument strings were distinct, reading db/queries/events.py, api/routers/sessions.py, api/middleware.py, 2256 and 2254 — the correct surface for the issue it is grounding. The cost shape comes from repeatedly re-fetching large files (db/queries/events.py exceeds 100KB) rather than from a loop.

Defect 1 — no per-child bound

design-pipeline v5 bounds MAX_ISSUES_PER_RUN = 2, MAX_DESIGN_ITERS = 2, MAX_DESIGN_LAPS = 4. Every one of those is a lap count. Nothing bounds an individual grounding child by wall-clock or spend, so a single child can run for hours at unbounded cost while every lap counter still reads healthy.

A lap bound does not bound cost when the expensive thing happens inside one lap.

Defect 2 — the watcher names a cause it cannot distinguish

The alarm's text prescribes: "the aios#2271 recurrence signature. CHECK the PR's required checks for one stuck in_progress — that is what watch_ci waits on forever. Then kill the child."

This run has no PR and no watch_ci call. It is a design sweep. Acting on that prescription would have destroyed roughly an hour of legitimate work.

At the watcher's altitude the two situations are identical: a suspended run with a long-lived child. The 2271 stuck-check case and an expensive-but-progressing grounding pass cannot be told apart from run state alone. Because the prescribed action is destructive, a confident wrong prescription is costly in a way a merely noisy alarm is not.

Suggested direction

For defect 1: a per-child spend and/or wall-clock ceiling that returns partial work rather than running unbounded. The ceiling should be derived from measured healthy grounding durations, not picked round.

For defect 2: have the alarm report the observation — child alive N minutes, $X spent, whether the run has a PR and whether a CI watch is present — and let the reader determine the mechanism. Do not address this by widening the burn threshold: that would blind the watcher to the real 2271 class it was built to catch.

Acceptance

  • A grounding child exceeding the ceiling terminates with partial output rather than continuing.
  • The bound is proven both ways: a normal-length child completes untouched; an over-budget child is stopped.
  • The alarm distinguishes "run has a stuck required check" from "run has no PR at all", and its text does not prescribe a remedy for a mechanism it did not verify.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    design-laps:2shovel-readyDesign settled, scope clear; ready to implement without further design discussion

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions