Skip to content

enrich(search,#11601): Search-11b-Deep-Part4 density 1122 -> 1588 — lectures chiffrees benchmark multi-seed (md-only) - #16546

Merged
jsboige merged 1 commit into
mainfrom
enrich/11601-search-11b-p4
Sep 18, 2026
Merged

jsboige merged 1 commit into
mainfrom
enrich/11601-search-11b-p4

Conversation

@jsboige

@jsboige jsboige commented Sep 17, 2026

Copy link
Copy Markdown
Owner

Grain: MED/notebook-dotnet — lane myia-po-2025:CoursIA — prev: MED/notebook-dotnet #16544

Enrichissement md-only de Search-11b-Metaheuristiques-Deep-Part4.ipynb (marathon metaheuristiques, tranche benchmark multi-seed dim=5). Densité 1122 -> 1588 chars/cellule (seuil 1200, marge +388).

Contenu ajouté (7 cellules markdown, +4198 chars)

Prose ancrée sur les sorties commitées des cellules [11] (tableau moy±std, seeds {1,7,42}) et [14] (taux de succès f<1e-3) :

  • Intro NFL : le benchmark comme instance locale du théorème — la question est (algorithme, paysage, dimension), pas « le meilleur algorithme ».
  • §1 Setup : les 3 invariants (dim 5, bornes $[-5,10]^5$, budget égal) + pourquoi multi-seed ; moyenne ET taux racontent des histoires différentes.
  • §2 Benchmarks : les 4 défauts distincts des paysages (Sphere témoin, Rastrigin pièges, Rosenbrock vallée, Ackley plateau).
  • §5 Tableau : comment lire moy±std ; trois lectures sautent aux yeux (PSO crevé à 3.6816 sur Rastrigin, Rosenbrock sans gagnant net — ABC 0.4434, SA jamais gagnant jamais catastrophique).
  • Lecture du tableau : le choc dimensionnel (PSO 0.0000 en dim 2 -> 3.6816 ± 0.9163 en dim 5) ; pourquoi Rosenbrock résiste à tous (vallée étroite, pas corrélés).
  • §6 Taux : lecture chiffrée (Rosenbrock 0% partout — ABC 440× au-dessus du seuil ; Rastrigin ABC 67% seul) ; moyenne ≠ taux (SA 0.0260/0% vs ABC 0.0034/67%).
  • §7 Exercices : les 3 questions ouvertes (dimension, 4e algorithme, équité du budget).

Aucun nouveau heading ; paragraphes append aux cellules existantes.

Validation

See #11601

…ectures chiffrees benchmark multi-seed (md-only)

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

⚠️ Prose/output review needed in the notebooks this PR changed: a numeric value is not anchored, an explicit relation is contradicted, or its evidence is missing. These cases remain distinct in the JSON report; the signal is advisory, NOT a merge gate.

Scope = notebooks CHANGED in this PR, not the whole corpus. Explicit claim-check relations resolve only against named CLAIM_METRICS from the local output window and are classified SUPPORTED, CONTRADICTED, or UNPROVEN.
The markdown-claims-output-report run artifact contains the structured JSON report. See python scripts/check_markdown_claims_output.py --help for re-running locally.
Detector rationale: c.290 / c.331 / PR #11435 numeric pathology, extended with low-noise relational evidence.

@github-actions

Copy link
Copy Markdown
Contributor

Notebook outputs-required (H.4 schema): PASS (every code cell carries an outputs: list)

@github-actions

Copy link
Copy Markdown
Contributor

Notebook PR Validation: PASS

  • Notebooks checked: 1
  • Code cells validated: 9
  • Result: All passed

Checks: H.1 (no errors), H.3 (execution_count), C.1 (no banned patterns)
Non-Python kernels (.NET/Lean): C.1 + errors only (execution_count advisory)
QuantConnect notebooks: C.1 + errors only (require QC Cloud for execution)

@github-actions

Copy link
Copy Markdown
Contributor

Golden-Set Execution (H.7 P3)

✅ 8/8 notebooks passed (certified reproducible)

Notebook Status Time
2.1-Workflow-ML.ipynb ✅ SUCCESS 3.2s
2.2-Descente-de-gradient.ipynb ✅ SUCCESS 4.0s
2.3-Regression-lineaire-logistique.ipynb ✅ SUCCESS 3.7s
2.4-Arbres-Forets-Ensembles.ipynb ✅ SUCCESS 4.9s
Search-01-StateSpace.ipynb ✅ SUCCESS 3.5s
SL-1-LogicalLearning.ipynb ✅ SUCCESS 2.2s
rl_4_multi_armed_bandits.ipynb ✅ SUCCESS 18.1s
GameTheory-04c-NashExistence-Python.ipynb ✅ SUCCESS 2.6s

Pinned lockfile: scripts/notebook_tools/golden_set.lock.txt (H.7 P3, axe A #4208)

@clusterManager-Myia clusterManager-Myia left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

VERDICT: LGTM

[Hermes] — review enrichissement Search-11b-Deep-Part4 (head 1095c72).

Ancres vérifiées dans le notebook au head : 3.6816 (3×), 0.4434 (4×), seuil 1e-3 (5×) — les trois lectures chiffrées de la prose (PSO Rastrigin, Rosenbrock ABC, taux f<1e-3) sont verbatim des sorties commises. Diff md-only : 7 cellules markdown ajoutées autour des cellules [11]/[14] existantes, aucune cellule code touchée. Les lectures qualitatives (Rastrigin pièges, Rosenbrock vallée, Ackley plateau) sont conceptuellement justes pour les 4 fonctions. 0 secret. Densité 1122→1588 confirmée par l'ampleur du diff.

[Hermes hermes-pr-review, cycle :15 17/09, host c92df397a786]

@myia-ai-01 myia-ai-01 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[ai-01 exact-head] APPROVED — 1095c727be8f8a346a607573c6b5edeec60a19a7

Body complet, 4 commentaires, 1 review, 0 thread et diff complet lus.

Les ajouts sont markdown-only et s'ancrent sur les deux sorties interprétées. Les pivots 3.6816, 0.4434, 0.0260, 0.0034, le seuil 1e-3 et les taux 0/67/100 % sont présents dans les outputs committés ou se redérivent directement ; Hermes les a re-vérifiés au même head. La distinction moyenne/taux, le choc dimension 2→5 et la résistance Rosenbrock sont formulés sans claim universel : le texte borne les conclusions à ce benchmark, ces trois seeds et cette dimension.

Les 9 cellules code et leurs outputs restent intacts ; validation, ratchets, positioning, claims/output et checks de domaine sont verts. B.0 clair, zéro thread. Le PR gate rouge actuel est traité fail-closed comme DWELL attendu depuis la tête ~15:09Z : cette approval n'autorise aucun merge avant maturité, réagrégation verte et dossier adjoint frais.

@jsboige

jsboige commented Sep 17, 2026

Copy link
Copy Markdown
Owner Author

[ADJOINT PREFLIGHT]
schema: 1
lane: myia-po-2025:CoursIA-2
pr: 16546
head: 1095c72
complete: true
body: read
comments-reviewed: 4
reviews-reviewed: 2
threads-reviewed: 0
threads-unresolved: 0
surfaces-sha256: 9e0dbcb6ddf77003d265162ec7b21cd6c320f45071cc262c229dd18b8d06a57a
diff-files: 1
diff-additions: 29
diff-deletions: 7
checks: latest-wins-green
b0: clear
scope: pass
domain: pass
verdict: READY
[/ADJOINT PREFLIGHT]

@jsboige

jsboige commented Sep 17, 2026

Copy link
Copy Markdown
Owner Author

[ADJOINT PREFLIGHT]
schema: 1
lane: myia-po-2025:CoursIA-2
pr: 16546
head: 1095c72
complete: true
body: read
comments-reviewed: 5
reviews-reviewed: 2
threads-reviewed: 0
threads-unresolved: 0
surfaces-sha256: 618c0393a4df4e19c3270f1a5b4533022b84662d16f18db2759a58e5e21e04d1
diff-files: 1
diff-additions: 29
diff-deletions: 7
checks: latest-wins-green
b0: clear
scope: pass
domain: pass
verdict: READY
[/ADJOINT PREFLIGHT]

@jsboige
jsboige merged commit cec0cb3 into main Sep 18, 2026
80 of 81 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants