diff --git a/MyIA.AI.Notebooks/GenAI/PostTraining/PT_15_controle_interpretabilite.ipynb b/MyIA.AI.Notebooks/GenAI/PostTraining/PT_15_controle_interpretabilite.ipynb index 8e23853d98..0d405c7e6e 100644 --- a/MyIA.AI.Notebooks/GenAI/PostTraining/PT_15_controle_interpretabilite.ipynb +++ b/MyIA.AI.Notebooks/GenAI/PostTraining/PT_15_controle_interpretabilite.ipynb @@ -941,10 +941,10 @@ "source": [ "## Références\n", "\n", - "- Arditi, A. et al. (2024). *Refusal in LLMs is mediated by a single direction.* arXiv:2406.11717.\n", + "- Arditi, A. et al. (2024). *Refusal in Language Models Is Mediated by a Single Direction.* arXiv:2406.11717.\n", "- Turner, A. et al. (2024). *Steering Language Models with Activation Engineering.* ai-alignment.com.\n", "- Lin, J. et al. (2024). *Universal jailbreak backdoors from poisoned human feedback.* / optimisation de la projection (discuté dans R14 §3.2.1b).\n", - "- Liu, J. et al. (2024). *Rethinking Machine Unlearning for Large Language Models.* arXiv:2402.08787.\n", + "- Liu, S. et al. (2024). *Rethinking Machine Unlearning for Large Language Models.* arXiv:2402.08787.\n", "- Farrell, M. et al. (2024). *Applying representation engineering to unlearning.* (discuté dans R14 §3.2.2).\n", "- Gade, P. et al. (2024) ; Lermen, A.-L. et al. (2024). *LoRA finetuning effectively undoes safety alignment.* (discutés dans R14 §3.2.2).\n", "- Jain, N. et al. (2024) ; Prakash, N. et al. (2024) ; Lee, B. X. et al. (2025). *Mechanistically analyzing the effects of finetuning on LLMs.* (R14 §3.2.2).\n", diff --git a/MyIA.AI.Notebooks/GenAI/PostTraining/PT_16_vericoding_formal_verification.ipynb b/MyIA.AI.Notebooks/GenAI/PostTraining/PT_16_vericoding_formal_verification.ipynb index cb0686902a..08abf0a1f3 100644 --- a/MyIA.AI.Notebooks/GenAI/PostTraining/PT_16_vericoding_formal_verification.ipynb +++ b/MyIA.AI.Notebooks/GenAI/PostTraining/PT_16_vericoding_formal_verification.ipynb @@ -1555,7 +1555,7 @@ "\n", "**Le vericoding en une phrase** : la preuve formelle rend la récompense *incontournable* — le modèle ne peut pas gagner sans satisfaire la spécification — mais la **qualité de la spécification devient le nouveau maillon faible**, et c'est un travail d'ingénieur, pas de vérificateur.\n", "\n", - "**Références** — Bursuc, Trimponas, Sato, Nikolić, *Vericoding: LLMs for Formally Verified Code Generation* (arXiv:2509.22908, 2025) ; benchmark : github.com/Beneficial-AI-Foundation/vericoding-benchmark (MIT) ; tâche LC0033 : benchmark CLEVER (fig. 8 du papier) ; gate du dépôt : workflow `lean-axiom` / `LeanVerifier.check_axioms`.\n", + "**Références** — Bursuc, S., Ehrenborg, T., Lin, S., et al. (13 auteurs, incl. M. Tegmark), *A Benchmark for Vericoding: Formally Verified Program Synthesis* (arXiv:2509.22908, 2025) ; benchmark : github.com/Beneficial-AI-Foundation/vericoding-benchmark (MIT) ; tâche LC0033 : benchmark CLEVER (fig. 8 du papier) ; gate du dépôt : workflow `lean-axiom` / `LeanVerifier.check_axioms`.\n", "\n", "**Séquence suivante** : PT-11b/PT-11c appliquent RLVR à des vérificateurs SymPy/Z3 — l'extension naturelle de ce notebook serait d'entraîner (GRPO) le petit modèle local *sur* la récompense Dafny, en bouclant la chaîne : génération → preuve → gradient.\n" ] diff --git a/MyIA.AI.Notebooks/IIT/ICT-Series/ICT-37-FLens-BeliefState.ipynb b/MyIA.AI.Notebooks/IIT/ICT-Series/ICT-37-FLens-BeliefState.ipynb index af865f3161..ea425663a2 100644 --- a/MyIA.AI.Notebooks/IIT/ICT-Series/ICT-37-FLens-BeliefState.ipynb +++ b/MyIA.AI.Notebooks/IIT/ICT-Series/ICT-37-FLens-BeliefState.ipynb @@ -357,7 +357,7 @@ "ground-truth est la distribution filtree `P(s_t | obs_0..t)` -- un point du\n", "2-simplexe, pas l'etat cache.\n", "\n", - "### RRXOR (Riechers & Crutchfield 2018, arXiv:1706.00883)\n", + "### RRXOR (Riechers & Crutchfield 2017, arXiv:1706.00883)\n", "Le processus repete les triplets `(r1, r2, r1 XOR r2)` : correlations par\n", "paires nulles, spectre plat, mais contrainte de triplet deterministe.\n", "Epsilon-machine a **5 etats causaux** (machine Mealy : emissions sur les\n", @@ -940,7 +940,7 @@ "\n", "- **arXiv:2405.15943** — Shai et al., *Transformers Represent Belief State Geometry in their Residual Stream* : papier fondateur du théorème de géométrie belief-state linéairement représentée dans le residual stream ; §2.2 définit la mise à jour `eta' = eta T^(x) / (eta T^(x) 1)` et §3.2 le RRXOR à 36 états de croyance.\n", "- **Marzen & Crutchfield 2017** — *Nearly maximally predictive features and their dimensions*, Phys. Rev. E 95(5):051301(R) : origine du processus Mess3 (correction d'attribution #16225 — l'ancienne mention « singh et al. 1994 » était erronée).\n", - "- **Riechers & Crutchfield 2018** — *Spectral Simplicity of Apparent Complexity, Part II* ([arXiv:1706.00883](https://arxiv.org/abs/1706.00883)) : définition du RRXOR (triplets r1, r2, r1 XOR r2), epsilon-machine à 5 états, S-MSP à 36 croyances (Fig. 4 et 7).\n", + "- **Riechers & Crutchfield 2017** — *Spectral Simplicity of Apparent Complexity, Part II* ([arXiv:1706.00883](https://arxiv.org/abs/1706.00883)) : définition du RRXOR (triplets r1, r2, r1 XOR r2), epsilon-machine à 5 états, S-MSP à 36 croyances (Fig. 4 et 7).\n", "- **arXiv:2602.02385** — *Transformers Learn Factored Representations* : base de la réimplémentation numpy-only (régimes orthogonal / probes linéaires).\n", "- **Dépôt ZM** — `Zeinab-Mohammadi/pytorch-AI-interpretability-transformer_ZM` : architecture minimale de référence, point de comparaison pour la migration future vers des activations réelles (cf. Epic #15475)." ] @@ -1050,4 +1050,4 @@ }, "nbformat": 4, "nbformat_minor": 5 -} \ No newline at end of file +} diff --git a/arxiv_attributions_registry.yaml b/arxiv_attributions_registry.yaml index 0b053fb146..66e8a97f99 100644 --- a/arxiv_attributions_registry.yaml +++ b/arxiv_attributions_registry.yaml @@ -83,12 +83,53 @@ attributions: source_pr: "#12838" correction: "Référence GNN+combinatorial: placeholder 'Graph Neural Networks for combinatorial optimization' (arxiv:2012.01806, lien mort) -> Cappart et al. 2021 (Combinatorial Optimization and Reasoning with Graph Neural Networks, arXiv:2102.09544, référence canonique)" date: "2026-08-24" - - # ===== Passe 5 — rescan du 2026-09-21 (PR #17296) ===== - - arxiv_id: "2301.05217" - notebook: "MyIA.AI.Notebooks/ML/DataScienceWithAgents/02-ML-Cours/2.9d-Features-Circulaires-Helice-Nombres.ipynb" - cell_index: 53 - expected_citation: "3. Nanda, N., Chan, L., et al. (2023). *Progress Measures for Grokking via Mechanistic Interpretability*. [arXiv:2301.05217](https://arxiv.org/abs/2301.05217)." - source_pr: "#17296" - correction: "Nanda et al. grokking: identifiant 2305.00493 (Basak 2023, physique nucleaire -- Estimation of collision centrality) -> 2301.05217 (Progress measures for grokking via mechanistic interpretability). Titre et auteurs etaient deja justes : seule l'attribution de l'ID etait fausse." - date: "2026-09-21" + + # ===== Passe 6 — rescan du 2026-09-24 (PR #17622) ===== + - arxiv_id: "1706.00883" + notebook: "MyIA.AI.Notebooks/IIT/ICT-Series/ICT-37-FLens-BeliefState.ipynb" + cell_index: 7 + expected_citation: "### RRXOR (Riechers & Crutchfield 2017, arXiv:1706.00883)" + source_pr: "#17622" + correction: "Riechers-Crutchiefield annee: 2018 -> 2017 (v1 2017/06, aucune ref journal sur le record arXiv)" + date: "2026-09-24" + + - arxiv_id: "1706.00883" + notebook: "MyIA.AI.Notebooks/IIT/ICT-Series/ICT-37-FLens-BeliefState.ipynb" + cell_index: 17 + expected_citation: "- **Riechers & Crutchfield 2017** — *Spectral Simplicity of Apparent Complexity, Part II*" + source_pr: "#17622" + correction: "Riechers-Crutchiefield annee: 2018 -> 2017 (cellule References, meme motif que cell 7)" + date: "2026-09-24" + + - arxiv_id: "2406.11717" + notebook: "MyIA.AI.Notebooks/GenAI/PostTraining/PT_15_controle_interpretabilite.ipynb" + cell_index: 20 + expected_citation: "- Arditi, A. et al. (2024). *Refusal in Language Models Is Mediated by a Single Direction.* arXiv:2406.11717." + source_pr: "#17622" + correction: "Arditi titre abrege 'Refusal in LLMs...' -> titre exact 'Refusal in Language Models Is Mediated by a Single Direction'" + date: "2026-09-24" + + - arxiv_id: "2402.08787" + notebook: "MyIA.AI.Notebooks/GenAI/PostTraining/PT_15_controle_interpretabilite.ipynb" + cell_index: 20 + expected_citation: "- Liu, S. et al. (2024). *Rethinking Machine Unlearning for Large Language Models.* arXiv:2402.08787." + source_pr: "#17622" + correction: "Initiale 1er auteur: Liu, J. -> Liu, S. (Liu, Sijia, mesure sur page abs)" + date: "2026-09-24" + + - arxiv_id: "2509.22908" + notebook: "MyIA.AI.Notebooks/GenAI/PostTraining/PT_16_vericoding_formal_verification.ipynb" + cell_index: 24 + expected_citation: "**Références** — Bursuc, S., Ehrenborg, T., Lin, S., et al. (13 auteurs, incl. M. Tegmark), *A Benchmark for Vericoding: Formally Verified Program Synthesis* (arXiv:2509.22908, 2025)" + source_pr: "#17622" + correction: "Titre et co-auteurs de prose inexistants ('Vericoding: LLMs for Formally Verified Code Generation' par Trimponas/Sato/Nikolic — aucun papier mesurable) -> vrai papier 'A Benchmark for Vericoding: Formally Verified Program Synthesis', Bursuc + 12 co-auteurs. Bursuc reste 1er auteur (garde anti-erreur-miroir #11163)" + date: "2026-09-24" + + # ===== Passe 5 — rescan du 2026-09-21 (PR #17296) ===== + - arxiv_id: "2301.05217" + notebook: "MyIA.AI.Notebooks/ML/DataScienceWithAgents/02-ML-Cours/2.9d-Features-Circulaires-Helice-Nombres.ipynb" + cell_index: 53 + expected_citation: "3. Nanda, N., Chan, L., et al. (2023). *Progress Measures for Grokking via Mechanistic Interpretability*. [arXiv:2301.05217](https://arxiv.org/abs/2301.05217)." + source_pr: "#17296" + correction: "Nanda et al. grokking: identifiant 2305.00493 (Basak 2023, physique nucleaire -- Estimation of collision centrality) -> 2301.05217 (Progress measures for grokking via mechanistic interpretability). Titre et auteurs etaient deja justes : seule l'attribution de l'ID etait fausse." + date: "2026-09-21" diff --git a/scripts/results/arxiv_rescan_2026-09-24.json b/scripts/results/arxiv_rescan_2026-09-24.json new file mode 100644 index 0000000000..db8b8cdae1 --- /dev/null +++ b/scripts/results/arxiv_rescan_2026-09-24.json @@ -0,0 +1,109 @@ +{ + "date": "2026-09-24", + "lane": "myia-po-2023:CoursIA", + "epic": "#11168", + "methode": "rescan repo-wide via scripts/notebook_tools/scan_arxiv_citations.py, diff contre l union des artefacts 2026-09-03 + 2026-09-21, verdicts par lecture des pages abs arxiv.org (meta tags citation_*) + confrontation a la prose", + "transport": { + "api_export_arxiv": "HTTP 406 intermittent (2/14 OK) malgre UA descriptif + mailto", + "substitut": "pages https://arxiv.org/abs/ (HTTP 200, metadonnees first-party via citation_title/citation_author/citation_date)", + "recherches_titre": "web (searxng) + pages abs versionnees (v1/v2) pour les mismatches apparents" + }, + "corpus": { + "notebooks_scannes": 1648, + "notebooks_citant": 145, + "ids_uniques": 165 + }, + "reconciliation": { + "ids_uniques_cites": 165, + "couverts_avant": 152, + "delta": 14, + "verifiee": true + }, + "delta": { + "ids": [ + "1706.00883", + "1710.05060", + "1803.03635", + "2006.10782", + "2202.01691", + "2206.06821", + "2402.05110", + "2402.08787", + "2406.11717", + "2501.16496", + "2509.22908", + "2510.18212", + "2608.14611", + "2609.11912" + ], + "verdicts": { + "ok": 10, + "drift_mineur": 3, + "mauvaise_attribution": 1, + "id_fantome": 0 + }, + "detail": { + "1706.00883": { + "verdict": "DRIFT-MINEUR", + "fix": "annee 2018 -> 2017 (v1 2017/06, aucune ref journal sur le record arXiv ; seules annees soutenues par l ID)", + "lieu": "ICT-37 cells 7+17" + }, + "1710.05060": { + "verdict": "OK", + "note": "Yudkowsky & Soares 2017, GT-04f" + }, + "1803.03635": { + "verdict": "OK", + "note": "annee prose 2019 = journal-ref ICLR 2019 presente sur le record arXiv lui-meme — auto-coherent, aucun edit (precedent Everitt #16389 vise une annee que l ID ne soutient pas ; ici le record la porte)" + }, + "2006.10782": { + "verdict": "OK", + "note": "Udrescu et al. AI Feynman 2.0, SL-14" + }, + "2202.01691": { + "verdict": "OK", + "note": "titre cite = titre EXACT de la v1 (2022/01/18), retitre en v2 le meme jour en Solving Dynamic Principal-Agent Problems — prose GT-04f cells 0+62 correcte (auteurs Mu/Zheng/Trott exacts), aucun edit. Caveat methodologique : comparer au titre COURANT de la page abs est structurellement aveugle aux retitrages de version — sur mismatch, verifier la version citee AVANT de corriger" + }, + "2206.06821": { + "verdict": "OK", + "note": "Blbaum/Gotz/Budhathoki/Mastakouri/Janzing DoWhy-GCM, GT-04f" + }, + "2402.05110": { + "verdict": "OK", + "note": "Michaud et al. Opening the AI black box, 2.9e" + }, + "2402.08787": { + "verdict": "DRIFT-MINEUR", + "fix": "initiale 1er auteur Liu, J. -> Liu, S. (Liu, Sijia, mesuree sur page abs)", + "lieu": "PT_15 cell 20" + }, + "2406.11717": { + "verdict": "DRIFT-MINEUR", + "fix": "titre abrege Refusal in LLMs... -> titre exact Refusal in Language Models Is Mediated by a Single Direction", + "lieu": "PT_15 cell 20" + }, + "2501.16496": { + "verdict": "OK", + "note": "Sharkey et al. Open Problems MechInterp, PT_15 + ICT-21" + }, + "2509.22908": { + "verdict": "MAUVAISE-ATTRIBUTION", + "fix": "titre de prose Vericoding: LLMs for Formally Verified Code Generation et co-auteurs Trimponas/Sato/Nikolic ne correspondent a AUCUN papier mesurable (recherches web seches) ; le vrai papier sous cet ID est A Benchmark for Vericoding: Formally Verified Program Synthesis, Bursuc + 12 co-auteurs (incl. Tegmark). Bursuc reste 1er auteur — garde anti-erreur-miroir #11163 respectee, pas de desattribution. Le contenu du notebook (benchmark 12504 taches, taux 82,2/26,8, CLEVER fig. 8) EST le benchmark — seule la ligne de reference etait fausse", + "lieu": "PT_16 cell 24" + }, + "2510.18212": { + "verdict": "OK", + "note": "Hendrycks/Song et al. Definition of AGI, 22b" + }, + "2608.14611": { + "verdict": "OK", + "note": "Casper et al. Singapore Consensus, PT_15 + 22b" + }, + "2609.11912": { + "verdict": "OK", + "note": "Becker/Greger/Peters Core in Approval-Based Committee Elections, 07-Committees-Core" + } + } + }, + "couverture_apres": 165 +}