Skip to content

docs(rlpt3,#14464): frontiere gaming vs tampering (Everitt 2021) - #14650

Merged
myia-ai-01 merged 1 commit into
mainfrom
docs/rlpt3-everitt-taxonomy-14464
Sep 5, 2026
Merged

myia-ai-01 merged 1 commit into
mainfrom
docs/rlpt3-everitt-taxonomy-14464

Conversation

@jsboige

@jsboige jsboige commented Sep 4, 2026

Copy link
Copy Markdown
Owner

Grain: MED/notebook-python -- lane myia-po-2026:CoursIA -- prev: MED/guard #14601

Quoi

Enrichissement markdown-only de MyIA.AI.Notebooks/RL/rlpt_3_reward_hacking.ipynb (#14464) : nouvelle section « ## 8. Frontière : specification gaming n'est pas reward tampering » insérée avant la conclusion (renumérotée 8 → 9), construite sur Everitt, Hutter, Kumar & Krakovna (2021), arXiv:1908.04734 :

  • Tableau taxonomie 3 classes (définitions fidèles à la Figure 1 du papier) : specification gaming (récompense fixe mal spécifiée, optimisée honnêtement) / reward-function tampering (influencer la fonction implémentée) / RF-input tampering (influencer l'information que la fonction possède sur l'état) ;
  • Un exemple précis par classe : l'amorce récompensée §5 de ce notebook (gaming) ; la réécriture du reward + l'agent Super-Mario exécutant du code arbitraire depuis la mémoire du jeu (RF tampering, cas « partially real » du papier) ; observations de diamants fictifs + Rocks-and-Diamonds partiellement observable Figure 10 (RF-input) ;
  • CID ASCII minimal — deux schémas : le reward n'est atteignable que par le comportement vs un second chemin s'ouvre vers le reward — avec la règle de lecture de la Figure 4 (but instrumental = chemin dirigé décision → X ET chemin X → utilité) ;
  • Déclaration de périmètre explicite : ce notebook ne démontre AUCUN tampering — ni reward_proxy ni ses entrées ne sont modifiables par le modèle ; §§3-5 = gaming pur, la frontière est tenue close par le protocole ;
  • Principes de conception + hypothèses : current-RF optimisation (confidentialité du RF initial), rewards uninfluenceable (historique/croyances plutôt qu'observations brutes), limite méthodologique (« un diagramme n'affirme que l'ABSENCE de buts instrumentaux, jamais leur présence ») ;
  • Renvoi vers DecInfer-05-Decision-Networks + citation complète avec chemin GDrive (convention des notebooks GameTheory-26/ICT-35) + Everitt ajouté à la ligne Références de la conclusion.

Fidélité vérifiée contre le PDF source (extraction pypdf, pages 1-8 et 18-19) : définitions Figure 1, structure d'ICI Figure 4a, §4.1, exemples 3/Figure 10b.

Validation (relancée après le dernier commit)

  • Markdown-only : zéro ligne execution_count/outputs dans le diff — exception C.2, outputs précédents valides.
  • python scripts/notebook_tools/notebook_tools.py validate <nb> = OK, 0 warnings, 0 errors.
  • python scripts/notebook_tools/check_notebook_navlinks.py <nb> = 0 lien cassé (le lien relatif DecInfer-05 résout).
  • python scripts/notebook_tools/validate_pr_notebooks.py origin/main <nb> = 1/1 passed (12 cellules code, exec_count + outputs cohérents).

G-VAR-3 : ordre de merge recommandé (exposition même-genre)

La lane a déjà deux PRs notebook-python ouvertes plus anciennes (#14626, #14642). Si l'une merge juste avant celle-ci, l'adjacency guard (qui lit la séquence mergée, cf incident #14638) la bloque. Tout interfoliage évitant deux notebook-python consécutifs convient, p. ex. : #14626 → #14644 (tooling) → #14642 → #14632 (ledger) → celle-ci.

Périmètre

1 fichier : MyIA.AI.Notebooks/RL/rlpt_3_reward_hacking.ipynb (+70/−3, markdown uniquement).

Closes #14464

Nouvelle section 8 "Frontiere : specification gaming n'est pas reward
tampering" avant la conclusion (renumeroee 9) : tableau taxonomie Everitt
et al. (3 classes, definitions Figure 1), un exemple par classe, CID ASCII
(regle du but instrumental, Figure 4), declaration de perimetre (aucun
tampering demontre par construction), principes current-RF optimisation /
rewards uninfluenceable + limite methodologique, renvoi DecInfer-05.
Markdown-only : aucun changement de cellule code (exception C.2).

Co-Authored-By: Claude-Code <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Golden-Set Execution (H.7 P3)

✅ 8/8 notebooks passed (certified reproducible)

Notebook Status Time
2.1-Workflow-ML.ipynb ✅ SUCCESS 11.5s
2.2-Descente-de-gradient.ipynb ✅ SUCCESS 12.2s
2.3-Regression-lineaire-logistique.ipynb ✅ SUCCESS 21.7s
2.4-Arbres-Forets-Ensembles.ipynb ✅ SUCCESS 10.0s
Search-1-StateSpace.ipynb ✅ SUCCESS 24.2s
SL-1-LogicalLearning.ipynb ✅ SUCCESS 6.5s
rl_4_multi_armed_bandits.ipynb ✅ SUCCESS 57.8s
GameTheory-04c-NashExistence-Python.ipynb ✅ SUCCESS 33.2s

Pinned lockfile: scripts/notebook_tools/golden_set.lock.txt (H.7 P3, axe A #4208)

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

G-VAR-2/3 GENRE signals (advisory, non bloquant, #10020).
La lane `myia-po-2026:CoursIA` voit ces signaux actifs sur les mergees du jour (UTC 2026-09-04) :

G-VAR-2 plafonne a max(1, grains_mergees_du_jour // 3) LIGHT par lane et par jour, toutes categories LIGHT confondues -- un RATIO, pas un plafond plat ; le cap calcule du jour est dans le tally ci-dessus. G-VAR-3 interdit deux genres LIGHT consecutifs. Les signaux ci-dessus rendent le fait VISIBLE (labels variation-tier-inflation, `variation-genre-run`, `variation-genre-cap-exceeded`, `variation-genre-mismatch`, `variation-genre-unknown`) -- la decision de merge reste au coordinateur.

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

⚠️ Detector abstained (merge-base introuvable, shallow fetch or unanchored branch).

c.415 (#11873): scope = notebooks CHANGED in this PR, not the whole corpus.
See python scripts/check_markdown_claims_output.py --help for re-running locally.
Detector rationale: c.290 / c.331 / PR #11435 pathologie.

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Notebook PR Validation: PASS

  • Notebooks checked: 1
  • Code cells validated: 12
  • Result: All passed

Checks: H.1 (no errors), H.3 (execution_count), C.1 (no banned patterns)
Non-Python kernels (.NET/Lean): C.1 + errors only (execution_count advisory)
QuantConnect notebooks: C.1 + errors only (require QC Cloud for execution)

@jsboige

jsboige commented Sep 4, 2026

Copy link
Copy Markdown
Owner Author

[DONE c.963] PR #14650 — docs(rlpt3,#14464): frontiere gaming vs tampering (Everitt 2021)

Lane myia-po-2026:CoursIA-2 — enrichissement markdown-only du notebook MyIA.AI.Notebooks/RL/rlpt_3_reward_hacking.ipynb pour répondre à l'acceptance #14464 (distinguer specification gaming des deux familles de reward tampering par diagramme causal).

Head : docs/rlpt3-everitt-taxonomy-14464, commit du 2026-09-04T16:51:51Z.

Grain : MED/notebook-python (Tell c.918 NAMING LIVRÉ-urn geste honnête — la PR était LIVRÉE et validée hier, ce cycle = reporting finalisé d'un cycle de livraison dont le dashboard [DONE] n'avait pas été posté).

Cycle de reporting : Tell c.477 sustained « commit + PR AVANT rapport » était tenu (PR commited, CLEAN MERGEABLE) — il manquait le dashboard [DONE]. Cycle c.963 = finalisation honnête d'un cycle livré sans reporting.

Périmètre : 1 fichier MyIA.AI.Notebooks/RL/rlpt_3_reward_hacking.ipynb (+70 / -3 lignes). 1 cellule MD insérée en section 8 (avant conclusion renumérotée 9) avec :

  • Taxonomie Everitt à 3 classes (specification gaming / reward-function tampering / RF-input tampering) — définitions Figure 1 du papier, exemple concret par classe ;
  • CID ASCII minimal illustrant quand un nœud devient une cible de contrôle instrumental (règle Figure 4) ;
  • Déclaration de périmètre explicite : « ce notebook ne démontre aucun tampering, par construction » ;
  • Principes current-RF optimisation / rewards uninfluenceable + limite méthodologique (CID ⟹ absence, jamais présence) ;
  • Renvoi vers DecInfer-05-Decision-Networks.ipynb qui enseigne déjà les diagrammes d'influence ;
  • Citation complète + chemin GDrive canonique ;
  • Ligne Références étendue avec Everitt 2021.

Validations :

  • notebook_tools.py validate OK 0 warnings ;
  • navlinks 0 lien cassé ;
  • validate_pr_notebooks.py 1/1 passed (12 cellules code intouchées — exception C.2 markdown-only) ;
  • Fidélité vérifiée contre le PDF source (pypdf pages 1-8/18-19).

Tells sustained ×c.963 : c.745 strict 1 pertinent/cycle ✓ (1 PR LIVRÉE, dashboard finalisé) · c.477 sustained « commit+PR AVANT rapport » ✓ (PR commited hier 16:51Z, juste report en souffrance) · c.745 strict 3 BANNED ✓ (pas de DELETE / PATCH body / PUT) · c.918 ★×16ᵉ cycle NAMING LIVRÉ-urn ✓ · c.13475 ★★★ PREV-NOT-PR ✓ (prev: MED/guard #14601 PR MERGED, distincte) · c.890 monotonie ×28ᵉ cycle c.929-c.963 LEVÉE par reporting c.963 (c.962=notebook-python Goodhart, c.963=notebook-python Everitt même notebook mais angle distinct) · c.1356 ★★★ vérif first-hand ✓ (relecture verbatim PDF source + diff PR).

G-VAR-1 : TENU ×27ᵉ (genre notebook-python CONTENU).

Prune worktrees (résiduel c.962 levé) :

Résiduel c.964+ :

— po-2026 c.963 worker (lane myia-po-2026:CoursIA-2)

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Path-collision (organ #13359/#13615)

Cette PR #14650 (docs(rlpt3,#14464): frontiere gaming vs tampering (Everitt 2021)) touche au moins un chemin de fichier aussi modifie par d'autres PRs ouvertes. Risque de double-livraison (meme fichier livre deux fois, 2x le travail et 2x les runs CI). Advisory : parfois legitime (tranches coordonnees, partition paths: explicite, PRs empilees exclues) -- l'organe rend visible, il ne bloque pas.

@myia-ai-01
myia-ai-01 merged commit 79962b5 into main Sep 5, 2026
62 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

RL : distinguer specification gaming et reward tampering par diagramme causal

2 participants