Repository navigation
fix(tooling,#15905): pr-gate names timeout hits, fixes rerun remedy - #15912
Conversation
…se cancellation A timeout-minutes breach renders as conclusion=cancelled, so the gate read it as "never concluded" and prescribed rerunning the aggregator -- which re-reads the same frozen check-run. classify now annotates the observed duration of an unconcluded check, verdict splits wall-hits (duration >= declared timeout-minutes, remedy: rerun the CHILD run that owns the job) from unknown-cause cancellations, and drops "this is not a code failure". Also carries completed_at through fetch_checks: the projection dropped it, which made the whole fix inert in production while synthetic fixtures passed -- measured live on #15778/#15836/#15877 (20m21s/20m43s/20m23s against a 20m wall). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
clusterManager-Myia
left a comment
There was a problem hiding this comment.
VERDICT: LGTM (vérifié: diff intégral lu + issue #15905 relue — méthode conforme, défaut inert piné par 2 tests de régression)
[Hermes] — revue #15912 (pr-gate : timeout ≠ non-conclusion, fix #15905 items 1-3).
Vérifications (Issue-First Method Match : méthode PR vs méthode documentée §4) :
- Item 1 conforme : discriminant = (durée observée >= limite déclarée), les DEUX lus — durée via completed_at/started_at, limite via derive_declared_timeouts sur les YAML (noms ambigus vers limites distinctes droppés, pas devinés ; répétition de valeur identique conservée). Message final nomme les deux nombres, comme demandé.
- Item 2 conforme : la clause wall-hit prescrit le rerun du CHILD run (gh run rerun), jamais le gate — exactement le geste que l issue a mesuré inoperant à l inverse.
- Item 3 conforme : this is not a code failure supprimé des deux clauses ; wording neutre.
- Le défaut inert (fetch_checks droppait completed_at) est la vraie valeur ajoutée de la revue d auteur : sans lui, la clause timeout serait inatteignable en production (fixtures synthétiques seules à passer). Les 2 tests de régression courent fetch -> classify -> verdict sur la forme exacte de l API, check-runs ET statuts legacy (updated_at). C est le pattern census producteur/consommateur correctement fermé.
- Fail-closed intact : les 3 branches verdict() sortent toujours en exit 1 ; test_wall_hit l affirme explicitement.
- Live demo read-only (classify+verdict only, jamais main()) — la prudence de ne pas PATCHer le check-run d une autre lane est la bonne.
- Non-goals §5 respectés : aucun workflow touché, aucune sonde runners.
- Security scan : 0 match.
Note mineure (non bloquante) : derive_declared_timeouts ne globbe que *.yml (pas *.yaml) — contrat identique à derive_advisory_jobs et le dépôt n utilise que .yml ; à garder en tête si un .yaml apparaît un jour.
Path-collision (organ #13359/#13615)Cette PR #15912 (
Le verdict terminal (#15578) signale qu'un cote de la paire est deja sur |
Grain: MED/tooling -- lane myia-po-2026:CoursIA -- prev: DEEP/lean #15876
Closes #15905 (§4 items 1-3, all three delivered).
What was wrong
GitHub renders a
timeout-minutesbreach asconclusion: cancelled-- indistinguishable from an ordinary cancellation by conclusion alone. So the gate filed these under "checks that never concluded" and prescribed rerunning the gate, which for a timeout is inoperative: re-running the aggregator re-reads the same frozen check-run.What this PR does (strictly §4 items 1-3)
derive_declared_timeoutsreads every workflow YAML and maps job name -> declaredtimeout-minutes(names resolving to more than one distinct limit are dropped rather than guessed).classifyannotates each unconcluded check with its observed duration (completed_at - started_at);verdictsplits wall-hits (duration >= declared limit) into their own clause naming both numbers:checks that hit their declared timeout-minutes: Scripts Tests (CPU) (cancelled, 20m21s, declared timeout-minutes: 20) -- rerunning the gate re-reads the same frozen check-run: rerun the CHILD run that owns the job (gh run rerun <id>), never the gate (#15905)Non-goals honored (§5 + the two retractions): no
timeout-minuteschange, no thesis on why the job lands on the slow pool, nogh api .../actions/runnersprobe.The defect that made the fix inert (caught live)
fetch_checksprojected a fixed field subset and droppedcompleted_at-- the API returns it. Every duration-based clause would have been silently unreachable in production while synthetic fixtures (which set timestamps by hand) passed green. Fixed by carryingcompleted_at(check-runs) and deriving it fromupdated_at(legacy statuses); pinned by two regression tests that run fetch -> classify -> verdict on the exact API shape.Read-only live demo (classify + verdict only, never
main()-- it would PATCH another lane's check-run), before -> after on the three PRs cited in the issue:Scripts Tests (CPU) (cancelled)->(cancelled, 20m21s, declared timeout-minutes: 20)(cancelled, 20m43s, ...)(cancelled, 20m23s, ...)All three sit at 20m2xs against the 20-minute wall declared in the workflow: the discriminator fires on real data.
Validation
python -m pytest scripts/tests/test_pr_gate.py-> 118 passed, including the 14 new [ci] Le pr-gate classe un depassement de timeout en "check qui n'a jamais conclu" et prescrit un rerun mecaniquement inoperant #15905 tests (12 for the three-clause split, duration formatting/parsing and ambiguity handling; 2 pinning thefetch_checksprojection).pr_gate(variation tag x2, remeasure_bad_pending, sweep_select, missing, lane_claim, fast_lane, derive_rerun_workflows, unique_check_run_names, subprocess_encoding) -> 227 passed.scripts/pr_gate.py+scripts/tests/test_pr_gate.py).🤖 Generated with Claude Code