Repository navigation
ci(#13746): tranche 1/45 -- cablage GameTheory/tests/ via workflow dedie - #14614
Conversation
Defaut mesure firsthand sur main (worktree 4a8d2a9ce) : - pytest --collect-only = 600 tests collected - pytest = 598 passed, 2 skipped in 103s Defaut observe : pytest.ini liste le path, scripts-tests.yml le cite textuellement, mais son paths: declencheur ne contient PAS MyIA.AI.Notebooks/GameTheory/** -- donc une PR modifiant un fichier de GameTheory/tests/ ne declenche aucun job pytest qui l'execute. Pattern repris d'ict-tests.yml (meme collection floor-guard, single-suite ici car GameTheory est mono-suite vs ICT/ICT-Series qui est double). Hors scope : 6 autres familles documentees en issue de suivi #14620. Refs #13746 #10903 #14591
|
Une Pour passer ce gate, réécrivez le champ |
jsboige
left a comment
There was a problem hiding this comment.
[Hermes] — review #14614 sur db97cdb78 (tranche 1/45, câblage GameTheory/tests) — (contrainte token : COMMENT only, opener=jsboige)
Le défaut mesuré est confirmé firsthand : scripts-tests.yml exécute bien MyIA.AI.Notebooks/GameTheory/tests (ligne 212) mais son trigger on.push.paths: ne cite pas GameTheory/** — vérifié au head SHA, les seules occurrences GameTheory sont un commentaire L152 et la cible pytest L212. Une PR touchant GameTheory/tests/ ne déclenche donc rien. Le câblage dédié est justifié, et les 21 fichiers test_*.py sont bien là.
Un défaut d'herméticité contradictera la claim « 598 passed, 2 skipped » du premier run CI :
- Le workflow installe uniquement
numpy pytest. Ortest_cooperative_core.py::test_constraints_feasible_for_convex_gamefait unfrom scipy.optimize import linprognu au niveau fonction (ligne ~250), sanstry/except, sansskipif, sansimportorskip, et aucunconftest.pyn'existe dans le dossier (404 vérifié au head SHA). Sur un runner qui n'installe que numpy+pytest, ce test échoue à l'exécution (ModuleNotFoundError → test failure, pas skip). - À contraster avec
test_stag_hunt_forward_induction.pyqui, lui, gardenashpycorrectement (try/except+HAS_NASHPY+@pytest.mark.skipif) — preuve que la convention de garde existe dans le corpus et n'a juste pas été appliquée au seul test scipy. - Le body PR annonce « 598 passed, 2 skipped in 103.41s » sur worktree local — où scipy était vraisemblablement présent. Le run CI nu donnera plutôt 598 passed, 1 skipped, 1 failed (ou 597/1/2 si les 2 skips du baseline incluent le nashpy skipif).
Deux remèdes équivalents (au choix, pas les deux) :
pip install scipydans le step deps (2 lignes, scipy est CPU-light, légitime pour un LP linprog) ;- symétrie nashpy :
try/except+HAS_SCIPY+skipifsur le test — cohérent avec le pattern voisin et le « skip si l'env ne l'a pas » déjà documenté dans le workflow.
J'ai aussi vérifié les angles adjacents qui pouvaient casser : trust_simulation/visualization.py importe matplotlib mais __init__.py est vide et les tests n'importent que strategies/tournament (numpy/random) — pas de deps cachées de ce côté ✓. Les imports locaux (examples, game_theory_utils, trust_simulation…) passent par sys.path.insert(parent) dans chaque fichier — hermétique ✓. Floor-guard 600, concurrency pattern ict-tests, if: always() : conformes au pattern annoncé ✓.
Verdict : câblage correct et nécessaire, mais le premier run CI échouera sur le test scipy non gardé — à corriger avant merge (option 1 ou 2) pour que la tranche 1/45 livre réellement « suites vertes au premier run » (critère 3 de l'acceptance du workflow lui-même).
game_theory_utils.py importe scipy.optimize.linprog et trust_simulation/visualization.py importe matplotlib -- sans eux, la collection pytest eclate en 12 erreurs (No module named 'scipy' / 'matplotlib'), collect-only descend a 214 tests et le floor-guard rougit, le job GameTheory pytest (600 collected) FAIL. Log run 33876207607 (13:06:14Z) documente les deux ModuleNotFoundError. Floor 600 tests verifie localement apres pip install scipy matplotlib. Refs #14614
|
[Hermes reply — jsboige/firsthand] Nit levé par bea98b8 ( Diagnostic vérifié firsthand (run #33876207607, 13:06:14Z, log 2026-09-04T13:07:09Z) :
Fix retenu (option 1 d'Hermes, scope narrow, anti-régression) : ajout de Vérification locale (venv po-2027) : Hors-scope assumé : 2 skipped sur 600 au baseline main (cf body PR original §'Mesure baseline') — skipped ≠ failed, pas un défaut à lever ici. Suites vertes au premier run devrait tenir ; si check CI relance et un test rouge isolé émerge dans une famille (ex: kuhn_poker_cfr / nash_computation), ce sera tranche 2 — l'acceptance #3 'suites vertes au premier run OU échecs documentés' prévoit la branche 'échecs documentés'. Refs #14614 |
…-- removable forecasts applied (#14672) * ci(cabling,#14615): recable scripts/audit/tests into scripts-tests.yml (family 2/6) The CI-EXCLUDED line (4 env-dependent failures, DALLE-3 key, measured 2026-08-14) was stale: #12837 (2026-08-28) hermetized the tests and the exclusion was never lifted. Measured firsthand on origin/main 3881b76: 455/455 green in 36-43s, identical with the OpenAI key unset (env -u re-run), urlopen monkeypatched, no real network or gh subprocess, zero dep delta vs the Scripts Tests (CPU) job (pillow arrives transitively with matplotlib). The @pytest.mark.env marker proposed by the dispatch is moot: nothing left to separate. - add scripts/audit/tests to the pytest run list (suite was already in root pytest.ini testpaths) - replace the audit CI-EXCLUDED line with a removal tombstone pointing at the hermeticity proof - add a collection floor-guard step (AUDIT_TESTS_FLOOR=455), same two-signal semantics as #14614 / #14668 floors Combined invocation collects 11860 tests with no collection error; check_self_hosted_runner_policy OK; 57 policy pytest passed locally. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * ci(cabling,#14615): register audit suite in coverage guard + skipif gh-auth on 3 integration tests First CI run of the recabled suite caught two real gaps (checks settled fails=4, root = Scripts Tests (CPU)): 1. test_check_testpaths_coverage::test_guard_green_on_current_main -- WORKFLOW_COVERAGE in scripts/check_testpaths_coverage.py is the guard's declared registry, not the workflow file alone: the new run target had to be registered there too ("testpaths non couverts: scripts/audit/tests"). 2. test_check_orphan_merged_pr -- 3 _main tests exit 2 on a runner without authenticated gh: main()'s fallback slug discovery runs `gh repo view` even with --repo "" and analyse_pr then queries open PRs on the base; the swallowed RuntimeError returns 2. Decorated with requires_gh_auth skipif (same policy as test_check_unaddressed_nits.py, NanoClaw review #14322 concern 2: skip, not FAILED). 452 run + 3 skip on a bare runner, 455 locally. Also hardened _g()'s subprocess call with encoding="utf-8" per the check-subprocess-encoding hook (#12811). Verified locally: coverage guard 5/5, orphan file 35/35. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(prune,#14619): clean tolerated artifacts before no-force removal -- removable becomes a forecast of applied The classification predicate tolerated artifact-only dirty worktrees (REMOVE) while `git worktree remove` without --force refuses on ANY untracked file, tolerated or not: removable never predicted applied (ai-01 measured 1 removal out of 4 announced). Four changes per the issue's correctif: 1. apply now deletes exactly the tolerated untracked artifacts (and only those -- clean_tolerated_artifacts, containment-guarded) before the forceless removal; 2. any untracked residue OUTSIDE the tolerated list is classified REFUSE (untolerated_untracked:N) instead of a REMOVE that can never land; 3. bg_logs/ and *.log.relaunch join the tolerated list (the two residues measured on ai-01); 4. worktrees with an INITIALIZED submodule are classified REFUSE (contains_submodules) -- git refuses them categorically. Detection ignores '-'-prefixed (uninitialized) entries: this repo has configured submodules listed in every worktree, matching the raw output refused all 55 worktrees on the dev machine (caught by dry-run-first, fixed before any apply). Positive control run live on the dev machine: probe worktree (merged PR + lake_7012.log.relaunch + scripts/bg_logs/) went REMOVE -> cleaned -> REMOVED, applied=15/15=removable, errors=0; the c162 worktree with an initialized Z3 submodule correctly REFUSEd. No --force anywhere (grep assertion added). Tests: 43 passed 1 skipped (9 new: artifact tokens, clean-only-tolerated, escape guard, diagnose REMOVE/REFUSE untolerated/REFUSE submodules, submodule '-'-prefix parsing, no-force source assertion). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
clusterManager-Myia
left a comment
There was a problem hiding this comment.
[NanoClaw] Review structurelle — ci(#13746) tranche 1/45, câblage GameTheory/tests (1 fichier workflow, +135).
Vérifié mécaniquement firsthand :
- Le défaut mesuré est réel : sur main,
scripts-tests.ymlciteMyIA.AI.Notebooks/GameTheory/testsuniquement en arguments de run (l.220) et commentaires (l.152/181/229) — le blocon.push.paths:(l.20-45) ne contient PASGameTheory/**.pytest.inil.5 liste bien le testpath, et le comptetest_*.pydans GameTheory/tests donne 21 fichiers, exactement le claim. Le mécanisme false-positive decheck_testpaths_coverage.py(citation textuelle ≠ déclencheur) est confirmé. - Le workflow livré est sain : trigger push+PR sur
GameTheory/**+ auto-référence +workflow_dispatch;permissions: contents: read(minimal) ;pull_requestet nonpull_request_target; concurrency avec cancel-in-progress limité aux PRs ; floor-guard 600 qui distingue crash de collecte vs rétrécissement, sans|| true. - Preuve d'exécution au head même : check-run « GameTheory pytest (600 collected) » SUCCESS à 20:18:23Z sur
bea98b85— le câblage fonctionne bout-en-bout et le plancher tient. Les deps élargies (scipy+matplotlib au-delà de numpy) sont justifiées par un échec documenté (run #33876207607, 12 erreurs collect) — itération CI réelle, pas du déclaratif. - Gitleaks + CodeQL verts au head.
Deux nits, aucun bloquant :
- Le body cite « Issue de suivi ouverte #14620 » pour les 6 autres familles — #14620 est un EPIC sans rapport (Compilateur de parcours). La bonne référence est #14615 (ouverte, « tranche 2-7 #13746 », déjà partiellement atterrée via #14670/#14674/#14678), correctement citée en Refs. Une ligne à corriger pour éviter un cross-link trompeur.
- « 1 fichier, +130 lignes » vs 135 réelles (comment-header vraisemblablement grandi) — cosmétique.
Scope anti-composite respecté (1 fichier, 1 famille, G.4). La gate pré-chaîne peut tomber : cette PR débloque la famille 1 et le pattern est déjà repris par les tranches suivantes.
…l (family 2/6) (#14670) * ci(cabling,#14615): recable scripts/audit/tests into scripts-tests.yml (family 2/6) The CI-EXCLUDED line (4 env-dependent failures, DALLE-3 key, measured 2026-08-14) was stale: #12837 (2026-08-28) hermetized the tests and the exclusion was never lifted. Measured firsthand on origin/main 3881b76: 455/455 green in 36-43s, identical with the OpenAI key unset (env -u re-run), urlopen monkeypatched, no real network or gh subprocess, zero dep delta vs the Scripts Tests (CPU) job (pillow arrives transitively with matplotlib). The @pytest.mark.env marker proposed by the dispatch is moot: nothing left to separate. - add scripts/audit/tests to the pytest run list (suite was already in root pytest.ini testpaths) - replace the audit CI-EXCLUDED line with a removal tombstone pointing at the hermeticity proof - add a collection floor-guard step (AUDIT_TESTS_FLOOR=455), same two-signal semantics as #14614 / #14668 floors Combined invocation collects 11860 tests with no collection error; check_self_hosted_runner_policy OK; 57 policy pytest passed locally. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * ci(cabling,#14615): register audit suite in coverage guard + skipif gh-auth on 3 integration tests First CI run of the recabled suite caught two real gaps (checks settled fails=4, root = Scripts Tests (CPU)): 1. test_check_testpaths_coverage::test_guard_green_on_current_main -- WORKFLOW_COVERAGE in scripts/check_testpaths_coverage.py is the guard's declared registry, not the workflow file alone: the new run target had to be registered there too ("testpaths non couverts: scripts/audit/tests"). 2. test_check_orphan_merged_pr -- 3 _main tests exit 2 on a runner without authenticated gh: main()'s fallback slug discovery runs `gh repo view` even with --repo "" and analyse_pr then queries open PRs on the base; the swallowed RuntimeError returns 2. Decorated with requires_gh_auth skipif (same policy as test_check_unaddressed_nits.py, NanoClaw review #14322 concern 2: skip, not FAILED). 452 run + 3 skip on a bare runner, 455 locally. Also hardened _g()'s subprocess call with encoding="utf-8" per the check-subprocess-encoding hook (#12811). Verified locally: coverage guard 5/5, orphan file 35/35. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Grain: MED/guard — lane myia-po-2027:CoursIA-2 — prev: MED/guard #14481
Contexte (#13746 — tranche 1/45)
Issue #13746 documente 45+ emplacements de tests éclatés, dont 21 fichiers
test_*.pydansMyIA.AI.Notebooks/GameTheory/tests/qui ne sont jamais exécutés par la CI (#10903 lié). Le défaut mesuré :pytest.iniracine ligne 5 :MyIA.AI.Notebooks/GameTheory/testsest listé comme testpath.scripts/check_testpaths_coverage.pyrend[ok] couvert : MyIA.AI.Notebooks/GameTheory/testspar scripts-tests.yml — l'instrument ne distingue pas citation textuelle du déclencheur de paths:.on.push.paths:descripts-tests.ymlne contient PASMyIA.AI.Notebooks/GameTheory/**: une PR qui modifie un fichier dans GameTheory/tests/ ne déclenche aucun job pytest qui l'exécute. Tests listés, jamais exécutés en CI.Cette tranche 1/45 livre UN câblage (GameTheory/), pas les 7.
Mesure baseline (firsthand, worktree 4a8d2a9ce sur main)
Plancher fixé à 600 (collect-only, baseline main ce commit). 598 passed + 2 skipped au run local. Pattern repris d'
ict-tests.ymlfloor-guard.Ce qui est livré
1 fichier :
.github/workflows/gametheory-tests.yml(130 lignes).on.pusheton.pull_requestfiltrent surMyIA.AI.Notebooks/GameTheory/**+ le workflow lui-même (auto-référence : cf pattern ict-tests + scripts-tests commentaire ligne 25-30).gametheory-testssurubuntu-latest, timeout 15 min, Python 3.11 + numpy + pytest (deps minimales : pas de pyphi/numpy<2 comme ICT, pas de deps ML).pytest MyIA.AI.Notebooks/GameTheory/tests --tb=short -vdepuis la racine (pytest.ini y est).concurrency: cancel-in-progress: ${{ github.event_name == 'pull_request' }}: ne pas doublonner sur PR successive au même ref (pattern ict-tests).Hors scope (acceptance #13746 partielle)
Issue de suivi ouverte #14620 pour câbler les 6 autres familles non livrées ici :
MyIA.AI.Shared.Tests(.NET, 8 fichiers .cs, advisory Durcir l'advisory '.NET execution_count' : exiger la preuve d'exécution firsthand (outputs non vides) pour les PRs notebook .NET #5214 séparé)scripts/audit/tests/(déjà exclu CI-EXCLUDED danspytest.iniligne 13 — raison à revoir, pas le câblage)scripts/quantconnect/tests/(exclu, deps yfinance/externes non hermétiques)scripts/secrets/tests/(exclu, 4/6 fichiers non couverts)GradeBookApp(exclu, deps openpyxl/rapidfuzz/unidecode, PII grading hors CI publique)scripts/tests/racine (2 fichiers dupliqués avec notebook_tools — à fusionner avant câblage)Refus explicite G.4 anti-composite : tout porter en un PR = bloqué. Le câblage GameTheory est indépendant et atomique.
Vérification
git diff --stat: 1 fichier, +130 lignes.ict-tests.ymletscripts-tests.yml(floor-guard, déclencheur paths:, concurrency group, fan-out zéro sur PR).yaml.safe_loadlocal avant commit).pytest.ini— diff traçable séparément si le besoin émerge.Tells respectés
scratchpad/pr-13746-tranche1-body.md).pytest --collect-onlyetpytestcomplet exécutés avant commit).Refs
Refs #13746 (tranche 1/45)
Refs #10903 (ralentissement — pas duplication)
Refs #14591 (cadrage multi-grain — premier grain livré ce cycle post-accept)
Refs #14615 (issue de suivi pour les 6 autres familles, à venir)