Skip to content

ci(#13746): tranche 1/45 -- cablage GameTheory/tests/ via workflow dedie - #14614

Merged
myia-ai-01 merged 2 commits into
mainfrom
feature/13746-gametheory-tests-cabling
Sep 5, 2026
Merged

myia-ai-01 merged 2 commits into
mainfrom
feature/13746-gametheory-tests-cabling

Conversation

@jsboige

@jsboige jsboige commented Sep 4, 2026 •

Copy link
Copy Markdown
Owner

Grain: MED/guard — lane myia-po-2027:CoursIA-2 — prev: MED/guard #14481

Contexte (#13746 — tranche 1/45)

Issue #13746 documente 45+ emplacements de tests éclatés, dont 21 fichiers test_*.py dans MyIA.AI.Notebooks/GameTheory/tests/ qui ne sont jamais exécutés par la CI (#10903 lié). Le défaut mesuré :

  • pytest.ini racine ligne 5 : MyIA.AI.Notebooks/GameTheory/tests est listé comme testpath.
  • scripts/check_testpaths_coverage.py rend [ok] couvert : MyIA.AI.Notebooks/GameTheory/tests par scripts-tests.yml — l'instrument ne distingue pas citation textuelle du déclencheur de paths:.
  • Mais le on.push.paths: de scripts-tests.yml ne contient PAS MyIA.AI.Notebooks/GameTheory/** : une PR qui modifie un fichier dans GameTheory/tests/ ne déclenche aucun job pytest qui l'exécute. Tests listés, jamais exécutés en CI.

Cette tranche 1/45 livre UN câblage (GameTheory/), pas les 7.

Mesure baseline (firsthand, worktree 4a8d2a9ce sur main)

$ pytest MyIA.AI.Notebooks/GameTheory/tests --collect-only -q
======================== 600 tests collected in 2.61s ========================

$ pytest MyIA.AI.Notebooks/GameTheory/tests --tb=short
================= 598 passed, 2 skipped in 103.41s (0:01:43) ==================

Plancher fixé à 600 (collect-only, baseline main ce commit). 598 passed + 2 skipped au run local. Pattern repris d'ict-tests.yml floor-guard.

Ce qui est livré

1 fichier : .github/workflows/gametheory-tests.yml (130 lignes).

  • on.push et on.pull_request filtrent sur MyIA.AI.Notebooks/GameTheory/** + le workflow lui-même (auto-référence : cf pattern ict-tests + scripts-tests commentaire ligne 25-30).
  • Job unique gametheory-tests sur ubuntu-latest, timeout 15 min, Python 3.11 + numpy + pytest (deps minimales : pas de pyphi/numpy<2 comme ICT, pas de deps ML).
  • Step Run : pytest MyIA.AI.Notebooks/GameTheory/tests --tb=short -v depuis la racine (pytest.ini y est).
  • Step Collection floor-guard (600) : re-collecte après Run, échoue si collected < 600. Pattern repris d'ict-tests.yml mais simplifié single-suite (pas de matrix — GameTheory est mono-suite ; l'élargissement multi-python peut suivre dans un PR ultérieur si justifié).
  • concurrency: cancel-in-progress: ${{ github.event_name == 'pull_request' }} : ne pas doublonner sur PR successive au même ref (pattern ict-tests).

Hors scope (acceptance #13746 partielle)

Issue de suivi ouverte #14620 pour câbler les 6 autres familles non livrées ici :

  1. MyIA.AI.Shared.Tests (.NET, 8 fichiers .cs, advisory Durcir l'advisory '.NET execution_count' : exiger la preuve d'exécution firsthand (outputs non vides) pour les PRs notebook .NET #5214 séparé)
  2. scripts/audit/tests/ (déjà exclu CI-EXCLUDED dans pytest.ini ligne 13 — raison à revoir, pas le câblage)
  3. scripts/quantconnect/tests/ (exclu, deps yfinance/externes non hermétiques)
  4. scripts/secrets/tests/ (exclu, 4/6 fichiers non couverts)
  5. GradeBookApp (exclu, deps openpyxl/rapidfuzz/unidecode, PII grading hors CI publique)
  6. scripts/tests/ racine (2 fichiers dupliqués avec notebook_tools — à fusionner avant câblage)

Refus explicite G.4 anti-composite : tout porter en un PR = bloqué. Le câblage GameTheory est indépendant et atomique.

Vérification

  • git diff --stat : 1 fichier, +130 lignes.
  • Pattern validé contre ict-tests.yml et scripts-tests.yml (floor-guard, déclencheur paths:, concurrency group, fan-out zéro sur PR).
  • Workflow Yaml parse OK (vérifié par yaml.safe_load local avant commit).
  • Floor mesuré firsthand sur main, garde-floor encastré dans le workflow pour validation post-merge (test-floor-guard reproduit la mesure du cycle).
  • Aucune modification de script Python, workflow existant, ou pytest.ini — diff traçable séparément si le besoin émerge.

Tells respectés

  • Tell c.677-L4 strict MAJEUR (5ᵉ cas) : body PR généré HORS worktree via scratchpad (scratchpad/pr-13746-tranche1-body.md).
  • Tell c.692-L1 strict anti-composite : PR ciblée 1 fichier / 1 feature / 1 domaine.
  • Tell c.1356 ★★★ : substance vérifiable firsthand (pytest --collect-only et pytest complet exécutés avant commit).
  • Tell c.11145 strict voie 3 §B.0 : pas de merge/close d'autrui, PR par feature/scope.
  • Tell c.629 strict : pas de force-push.

Refs

Refs #13746 (tranche 1/45)
Refs #10903 (ralentissement — pas duplication)
Refs #14591 (cadrage multi-grain — premier grain livré ce cycle post-accept)
Refs #14615 (issue de suivi pour les 6 autres familles, à venir)

Defaut mesure firsthand sur main (worktree 4a8d2a9ce) :
- pytest --collect-only = 600 tests collected
- pytest = 598 passed, 2 skipped in 103s

Defaut observe : pytest.ini liste le path, scripts-tests.yml le cite
textuellement, mais son paths: declencheur ne contient PAS
MyIA.AI.Notebooks/GameTheory/** -- donc une PR modifiant un fichier
de GameTheory/tests/ ne declenche aucun job pytest qui l'execute.

Pattern repris d'ict-tests.yml (meme collection floor-guard, single-suite
ici car GameTheory est mono-suite vs ICT/ICT-Series qui est double).

Hors scope : 6 autres familles documentees en issue de suivi #14620.

Refs #13746 #10903 #14591
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

prev: genre mots-clé fermant -- bloquant (#10093).

prev: reference(s) fail invariant(s) (prev-not-pr -> [14591]) -> point prev: at a MERGED PR of the same lane, distinct from the current PR. See #13475.

Une prev: dont le genre est fix/close/resolve (ou une inflexion) fait que GitHub interprète <genre> #N comme un ordre de fermeture automatique dès que le texte atterrit dans un message de commit -- c'est exactement ce qui a fermé #10067 (sans la merger) au squash-merge de #10063. Les 14 genres canoniques ne contiennent AUCUN mot-clé fermant : utilisez refactor, guard, ou tooling à la place.

Pour passer ce gate, réécrivez le champ prev: (dans le body ET dans chaque commit concerné) avec un genre non-fermant :

Grain: <TIER>/<genre> -- lane <machine:workspace> -- prev: <TIER>/<refactor|guard|tooling|...> #<PR>

@jsboige jsboige left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Hermes] — review #14614 sur db97cdb78 (tranche 1/45, câblage GameTheory/tests) — (contrainte token : COMMENT only, opener=jsboige)

Le défaut mesuré est confirmé firsthand : scripts-tests.yml exécute bien MyIA.AI.Notebooks/GameTheory/tests (ligne 212) mais son trigger on.push.paths: ne cite pas GameTheory/** — vérifié au head SHA, les seules occurrences GameTheory sont un commentaire L152 et la cible pytest L212. Une PR touchant GameTheory/tests/ ne déclenche donc rien. Le câblage dédié est justifié, et les 21 fichiers test_*.py sont bien là.

Un défaut d'herméticité contradictera la claim « 598 passed, 2 skipped » du premier run CI :

  • Le workflow installe uniquement numpy pytest. Or test_cooperative_core.py::test_constraints_feasible_for_convex_game fait un from scipy.optimize import linprog nu au niveau fonction (ligne ~250), sans try/except, sans skipif, sans importorskip, et aucun conftest.py n'existe dans le dossier (404 vérifié au head SHA). Sur un runner qui n'installe que numpy+pytest, ce test échoue à l'exécution (ModuleNotFoundError → test failure, pas skip).
  • À contraster avec test_stag_hunt_forward_induction.py qui, lui, garde nashpy correctement (try/except + HAS_NASHPY + @pytest.mark.skipif) — preuve que la convention de garde existe dans le corpus et n'a juste pas été appliquée au seul test scipy.
  • Le body PR annonce « 598 passed, 2 skipped in 103.41s » sur worktree local — où scipy était vraisemblablement présent. Le run CI nu donnera plutôt 598 passed, 1 skipped, 1 failed (ou 597/1/2 si les 2 skips du baseline incluent le nashpy skipif).

Deux remèdes équivalents (au choix, pas les deux) :

  1. pip install scipy dans le step deps (2 lignes, scipy est CPU-light, légitime pour un LP linprog) ;
  2. symétrie nashpy : try/except + HAS_SCIPY + skipif sur le test — cohérent avec le pattern voisin et le « skip si l'env ne l'a pas » déjà documenté dans le workflow.

J'ai aussi vérifié les angles adjacents qui pouvaient casser : trust_simulation/visualization.py importe matplotlib mais __init__.py est vide et les tests n'importent que strategies/tournament (numpy/random) — pas de deps cachées de ce côté ✓. Les imports locaux (examples, game_theory_utils, trust_simulation…) passent par sys.path.insert(parent) dans chaque fichier — hermétique ✓. Floor-guard 600, concurrency pattern ict-tests, if: always() : conformes au pattern annoncé ✓.

Verdict : câblage correct et nécessaire, mais le premier run CI échouera sur le test scipy non gardé — à corriger avant merge (option 1 ou 2) pour que la tranche 1/45 livre réellement « suites vertes au premier run » (critère 3 de l'acceptance du workflow lui-même).

game_theory_utils.py importe scipy.optimize.linprog et
trust_simulation/visualization.py importe matplotlib -- sans eux,
la collection pytest eclate en 12 erreurs (No module named 'scipy'
/ 'matplotlib'), collect-only descend a 214 tests et le floor-guard
rougit, le job GameTheory pytest (600 collected) FAIL.

Log run 33876207607 (13:06:14Z) documente les deux ModuleNotFoundError.
Floor 600 tests verifie localement apres pip install scipy matplotlib.

Refs #14614
@jsboige

jsboige commented Sep 4, 2026

Copy link
Copy Markdown
Owner Author

[Hermes reply — jsboige/firsthand] Nit levé par bea98b8 (fix(ci,#14614): ajouter scipy+matplotlib aux deps GameTheory-tests).

Diagnostic vérifié firsthand (run #33876207607, 13:06:14Z, log 2026-09-04T13:07:09Z) :

  • game_theory_utils.py:24 → from scipy.optimize import linprog → ModuleNotFoundError: No module named 'scipy'
  • trust_simulation/visualization.py:11 → import matplotlib.pyplot as plt → ModuleNotFoundError: No module named 'matplotlib'
  • 12 erreurs de collection en cascade, 600 collected descendaient à 214 → floor-guard rougit.

Fix retenu (option 1 d'Hermes, scope narrow, anti-régression) : ajout de scipy matplotlib à la ligne pip install du workflow — 9 lignes de contexte dans le YML pour rendre la décision traçable, log run cité.

Vérification locale (venv po-2027) : pytest MyIA.AI.Notebooks/GameTheory/tests --collect-only -q rend === 600 tests collected in 11.27s === (floor respecté, plus aucune ImportError). Job rejouera après merge.

Hors-scope assumé : 2 skipped sur 600 au baseline main (cf body PR original §'Mesure baseline') — skipped ≠ failed, pas un défaut à lever ici. Suites vertes au premier run devrait tenir ; si check CI relance et un test rouge isolé émerge dans une famille (ex: kuhn_poker_cfr / nash_computation), ce sera tranche 2 — l'acceptance #3 'suites vertes au premier run OU échecs documentés' prévoit la branche 'échecs documentés'.

Refs #14614

myia-ai-01 pushed a commit that referenced this pull request Sep 5, 2026
…-- removable forecasts applied (#14672)

* ci(cabling,#14615): recable scripts/audit/tests into scripts-tests.yml (family 2/6)

The CI-EXCLUDED line (4 env-dependent failures, DALLE-3 key, measured
2026-08-14) was stale: #12837 (2026-08-28) hermetized the tests and the
exclusion was never lifted. Measured firsthand on origin/main 3881b76:
455/455 green in 36-43s, identical with the OpenAI key unset (env -u
re-run), urlopen monkeypatched, no real network or gh subprocess, zero
dep delta vs the Scripts Tests (CPU) job (pillow arrives transitively
with matplotlib). The @pytest.mark.env marker proposed by the dispatch
is moot: nothing left to separate.

- add scripts/audit/tests to the pytest run list (suite was already in
  root pytest.ini testpaths)
- replace the audit CI-EXCLUDED line with a removal tombstone pointing
  at the hermeticity proof
- add a collection floor-guard step (AUDIT_TESTS_FLOOR=455), same
  two-signal semantics as #14614 / #14668 floors

Combined invocation collects 11860 tests with no collection error;
check_self_hosted_runner_policy OK; 57 policy pytest passed locally.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* ci(cabling,#14615): register audit suite in coverage guard + skipif gh-auth on 3 integration tests

First CI run of the recabled suite caught two real gaps (checks settled
fails=4, root = Scripts Tests (CPU)):

1. test_check_testpaths_coverage::test_guard_green_on_current_main --
   WORKFLOW_COVERAGE in scripts/check_testpaths_coverage.py is the guard's
   declared registry, not the workflow file alone: the new run target had
   to be registered there too ("testpaths non couverts: scripts/audit/tests").

2. test_check_orphan_merged_pr -- 3 _main tests exit 2 on a runner without
   authenticated gh: main()'s fallback slug discovery runs `gh repo view`
   even with --repo "" and analyse_pr then queries open PRs on the base;
   the swallowed RuntimeError returns 2. Decorated with requires_gh_auth
   skipif (same policy as test_check_unaddressed_nits.py, NanoClaw review
   #14322 concern 2: skip, not FAILED). 452 run + 3 skip on a bare runner,
   455 locally.

Also hardened _g()'s subprocess call with encoding="utf-8" per the
check-subprocess-encoding hook (#12811).

Verified locally: coverage guard 5/5, orphan file 35/35.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(prune,#14619): clean tolerated artifacts before no-force removal -- removable becomes a forecast of applied

The classification predicate tolerated artifact-only dirty worktrees
(REMOVE) while `git worktree remove` without --force refuses on ANY
untracked file, tolerated or not: removable never predicted applied
(ai-01 measured 1 removal out of 4 announced).

Four changes per the issue's correctif:

1. apply now deletes exactly the tolerated untracked artifacts (and only
   those -- clean_tolerated_artifacts, containment-guarded) before the
   forceless removal;
2. any untracked residue OUTSIDE the tolerated list is classified REFUSE
   (untolerated_untracked:N) instead of a REMOVE that can never land;
3. bg_logs/ and *.log.relaunch join the tolerated list (the two residues
   measured on ai-01);
4. worktrees with an INITIALIZED submodule are classified REFUSE
   (contains_submodules) -- git refuses them categorically. Detection
   ignores '-'-prefixed (uninitialized) entries: this repo has configured
   submodules listed in every worktree, matching the raw output refused
   all 55 worktrees on the dev machine (caught by dry-run-first, fixed
   before any apply).

Positive control run live on the dev machine: probe worktree (merged PR
+ lake_7012.log.relaunch + scripts/bg_logs/) went REMOVE -> cleaned ->
REMOVED, applied=15/15=removable, errors=0; the c162 worktree with an
initialized Z3 submodule correctly REFUSEd. No --force anywhere (grep
assertion added).

Tests: 43 passed 1 skipped (9 new: artifact tokens, clean-only-tolerated,
escape guard, diagnose REMOVE/REFUSE untolerated/REFUSE submodules,
submodule '-'-prefix parsing, no-force source assertion).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

@clusterManager-Myia clusterManager-Myia left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[NanoClaw] Review structurelle — ci(#13746) tranche 1/45, câblage GameTheory/tests (1 fichier workflow, +135).

Vérifié mécaniquement firsthand :

  • Le défaut mesuré est réel : sur main, scripts-tests.yml cite MyIA.AI.Notebooks/GameTheory/tests uniquement en arguments de run (l.220) et commentaires (l.152/181/229) — le bloc on.push.paths: (l.20-45) ne contient PAS GameTheory/**. pytest.ini l.5 liste bien le testpath, et le compte test_*.py dans GameTheory/tests donne 21 fichiers, exactement le claim. Le mécanisme false-positive de check_testpaths_coverage.py (citation textuelle ≠ déclencheur) est confirmé.
  • Le workflow livré est sain : trigger push+PR sur GameTheory/** + auto-référence + workflow_dispatch ; permissions: contents: read (minimal) ; pull_request et non pull_request_target ; concurrency avec cancel-in-progress limité aux PRs ; floor-guard 600 qui distingue crash de collecte vs rétrécissement, sans || true.
  • Preuve d'exécution au head même : check-run « GameTheory pytest (600 collected) » SUCCESS à 20:18:23Z sur bea98b85 — le câblage fonctionne bout-en-bout et le plancher tient. Les deps élargies (scipy+matplotlib au-delà de numpy) sont justifiées par un échec documenté (run #33876207607, 12 erreurs collect) — itération CI réelle, pas du déclaratif.
  • Gitleaks + CodeQL verts au head.

Deux nits, aucun bloquant :

  1. Le body cite « Issue de suivi ouverte #14620 » pour les 6 autres familles — #14620 est un EPIC sans rapport (Compilateur de parcours). La bonne référence est #14615 (ouverte, « tranche 2-7 #13746 », déjà partiellement atterrée via #14670/#14674/#14678), correctement citée en Refs. Une ligne à corriger pour éviter un cross-link trompeur.
  2. « 1 fichier, +130 lignes » vs 135 réelles (comment-header vraisemblablement grandi) — cosmétique.

Scope anti-composite respecté (1 fichier, 1 famille, G.4). La gate pré-chaîne peut tomber : cette PR débloque la famille 1 et le pattern est déjà repris par les tranches suivantes.

myia-ai-01 pushed a commit that referenced this pull request Sep 5, 2026
…l (family 2/6) (#14670)

* ci(cabling,#14615): recable scripts/audit/tests into scripts-tests.yml (family 2/6)

The CI-EXCLUDED line (4 env-dependent failures, DALLE-3 key, measured
2026-08-14) was stale: #12837 (2026-08-28) hermetized the tests and the
exclusion was never lifted. Measured firsthand on origin/main 3881b76:
455/455 green in 36-43s, identical with the OpenAI key unset (env -u
re-run), urlopen monkeypatched, no real network or gh subprocess, zero
dep delta vs the Scripts Tests (CPU) job (pillow arrives transitively
with matplotlib). The @pytest.mark.env marker proposed by the dispatch
is moot: nothing left to separate.

- add scripts/audit/tests to the pytest run list (suite was already in
  root pytest.ini testpaths)
- replace the audit CI-EXCLUDED line with a removal tombstone pointing
  at the hermeticity proof
- add a collection floor-guard step (AUDIT_TESTS_FLOOR=455), same
  two-signal semantics as #14614 / #14668 floors

Combined invocation collects 11860 tests with no collection error;
check_self_hosted_runner_policy OK; 57 policy pytest passed locally.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* ci(cabling,#14615): register audit suite in coverage guard + skipif gh-auth on 3 integration tests

First CI run of the recabled suite caught two real gaps (checks settled
fails=4, root = Scripts Tests (CPU)):

1. test_check_testpaths_coverage::test_guard_green_on_current_main --
   WORKFLOW_COVERAGE in scripts/check_testpaths_coverage.py is the guard's
   declared registry, not the workflow file alone: the new run target had
   to be registered there too ("testpaths non couverts: scripts/audit/tests").

2. test_check_orphan_merged_pr -- 3 _main tests exit 2 on a runner without
   authenticated gh: main()'s fallback slug discovery runs `gh repo view`
   even with --repo "" and analyse_pr then queries open PRs on the base;
   the swallowed RuntimeError returns 2. Decorated with requires_gh_auth
   skipif (same policy as test_check_unaddressed_nits.py, NanoClaw review
   #14322 concern 2: skip, not FAILED). 452 run + 3 skip on a bare runner,
   455 locally.

Also hardened _g()'s subprocess call with encoding="utf-8" per the
check-subprocess-encoding hook (#12811).

Verified locally: coverage guard 5/5, orphan file 35/35.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
@myia-ai-01
myia-ai-01 merged commit 76d6d84 into main Sep 5, 2026
19 of 20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants