Observation
lane-child-burn-watch fired on run wfr_01M1WEDVR1E7RRZBAX5NZ7PC3A (workflow design-pipeline v5, trigger design-sweep-15m) reporting a suspended run with a live child.
Measured directly against the usage ledger, not run state:
| time |
child events |
child cost |
| 23:41Z |
3,780 |
$35.04 |
| 23:56Z |
5,885 |
$56.42 |
That is roughly $85/hour and ~2,100 events per 15 minutes, for a grounding pass whose declared output schema is five small arrays (facts, seam_files, substrate_core, half_built_prior_art, thread_directives).
The child is not wedged. Across ~30 sampled tool calls, 22 argument strings were distinct, reading db/queries/events.py, api/routers/sessions.py, api/middleware.py, 2256 and 2254 — the correct surface for the issue it is grounding. The cost shape comes from repeatedly re-fetching large files (db/queries/events.py exceeds 100KB) rather than from a loop.
Defect 1 — no per-child bound
design-pipeline v5 bounds MAX_ISSUES_PER_RUN = 2, MAX_DESIGN_ITERS = 2, MAX_DESIGN_LAPS = 4. Every one of those is a lap count. Nothing bounds an individual grounding child by wall-clock or spend, so a single child can run for hours at unbounded cost while every lap counter still reads healthy.
A lap bound does not bound cost when the expensive thing happens inside one lap.
Defect 2 — the watcher names a cause it cannot distinguish
The alarm's text prescribes: "the aios#2271 recurrence signature. CHECK the PR's required checks for one stuck in_progress — that is what watch_ci waits on forever. Then kill the child."
This run has no PR and no watch_ci call. It is a design sweep. Acting on that prescription would have destroyed roughly an hour of legitimate work.
At the watcher's altitude the two situations are identical: a suspended run with a long-lived child. The 2271 stuck-check case and an expensive-but-progressing grounding pass cannot be told apart from run state alone. Because the prescribed action is destructive, a confident wrong prescription is costly in a way a merely noisy alarm is not.
Suggested direction
For defect 1: a per-child spend and/or wall-clock ceiling that returns partial work rather than running unbounded. The ceiling should be derived from measured healthy grounding durations, not picked round.
For defect 2: have the alarm report the observation — child alive N minutes, $X spent, whether the run has a PR and whether a CI watch is present — and let the reader determine the mechanism. Do not address this by widening the burn threshold: that would blind the watcher to the real 2271 class it was built to catch.
Acceptance
- A grounding child exceeding the ceiling terminates with partial output rather than continuing.
- The bound is proven both ways: a normal-length child completes untouched; an over-budget child is stopped.
- The alarm distinguishes "run has a stuck required check" from "run has no PR at all", and its text does not prescribe a remedy for a mechanism it did not verify.
Observation
lane-child-burn-watchfired on runwfr_01M1WEDVR1E7RRZBAX5NZ7PC3A(workflowdesign-pipelinev5, triggerdesign-sweep-15m) reporting a suspended run with a live child.Measured directly against the usage ledger, not run state:
That is roughly $85/hour and ~2,100 events per 15 minutes, for a grounding pass whose declared output schema is five small arrays (
facts,seam_files,substrate_core,half_built_prior_art,thread_directives).The child is not wedged. Across ~30 sampled tool calls, 22 argument strings were distinct, reading
db/queries/events.py,api/routers/sessions.py,api/middleware.py, 2256 and 2254 — the correct surface for the issue it is grounding. The cost shape comes from repeatedly re-fetching large files (db/queries/events.pyexceeds 100KB) rather than from a loop.Defect 1 — no per-child bound
design-pipelinev5 boundsMAX_ISSUES_PER_RUN = 2,MAX_DESIGN_ITERS = 2,MAX_DESIGN_LAPS = 4. Every one of those is a lap count. Nothing bounds an individual grounding child by wall-clock or spend, so a single child can run for hours at unbounded cost while every lap counter still reads healthy.A lap bound does not bound cost when the expensive thing happens inside one lap.
Defect 2 — the watcher names a cause it cannot distinguish
The alarm's text prescribes: "the aios#2271 recurrence signature. CHECK the PR's required checks for one stuck in_progress — that is what watch_ci waits on forever. Then kill the child."
This run has no PR and no
watch_cicall. It is a design sweep. Acting on that prescription would have destroyed roughly an hour of legitimate work.At the watcher's altitude the two situations are identical: a suspended run with a long-lived child. The 2271 stuck-check case and an expensive-but-progressing grounding pass cannot be told apart from run state alone. Because the prescribed action is destructive, a confident wrong prescription is costly in a way a merely noisy alarm is not.
Suggested direction
For defect 1: a per-child spend and/or wall-clock ceiling that returns partial work rather than running unbounded. The ceiling should be derived from measured healthy grounding durations, not picked round.
For defect 2: have the alarm report the observation — child alive N minutes, $X spent, whether the run has a PR and whether a CI watch is present — and let the reader determine the mechanism. Do not address this by widening the burn threshold: that would blind the watcher to the real 2271 class it was built to catch.
Acceptance