feat(post-training,#12653): PT-11c GRPO-RLVR cran 1.7B/2B structure notebook - #12885
Conversation
…otebook Grain CONTENU frais c.556 -> c.557 : creation structurelle de PT-11c sur base PT-11b, avec adaptations Qwen3-1.7B/2B (cran au-dessus 0.8B) : - Architecture compatible trl.GRPOTrainer + peft 0.20.0 (post-install c.556) - 4 seeds x 100 steps (vs 2x20 PT-11b) -> intra-seed DM informatif >=22 pts - Verdict memoire cible RTX 3080 Ti 16GB (pic VRAM mesure) - Verifier SymPy Tier-1 + Z3 N-queens Tier-2 repris verbatim PT-11b - rewardspy.watch_trl wrappe en fallback no-op si absent (env po-2026) - LOAD_MODEL_AND_TRAIN=False par defaut (mode lecture, code C.1-compliant) - 3 exercices C.1 : parser etendu sci, dataset geometrie, cross-cran 11b/11c - 33 cellules (18 md + 15 code), execution end-to-end 0 erreur Acceptance #12653 : - (1) structure notebook cree : OK ce cycle - (2) run live 4x100 steps : c.558 - (3) verdict honnete BEATS/NO BEATS/MECANISME_REPRO/INCONCLUSIVE + PR : c.559 README PostTraining MAJ : pedagogical_count 14->15, breakdown PostTraining=15, maturite BETA=10 + ALPHA=5, ligne PT-11c inseree entre PT-11b multi-seed et PT-12. Notebook commite avec outputs reels (C.2) - C.1 sans erreur volontaire. Co-Authored-By: Claude Haiku 4.5 (1M context) <noreply@anthropic.com>
jsboige
left a comment
There was a problem hiding this comment.
[Hermes] — review notebook-tier (structure PT-11c), vérifications firsthand sur le diff.
- Authenticité des outputs (règle C.2) : 15/15 cellules code exécutées,
execution_count1→15 contigu, 0 null. Les streams sont un vrai env probe : Python 3.11.9/Windows, torch 2.13.0+cpu,CUDA dispo : False, peft 0.20.0, trl 1.10.0, z3 4.16.0, sympy 1.14.0, etrewardspy : NOT INSTALLED (no-op wrapper...)— chaque claim d'env du body a son pendant littéral dans les outputs. Structure-only assumé : aucun résultat de training fabriqué (run live = c.558, hors scope déclaré). - Règle C.1 : 0
NotImplementedError/assert False/1/0; les 3 exos étudiants sont des stubsreturn None+TODO etudiantconformes. - Comptes : 33 cellules (18 md + 15 code) confirmées ; README
pedagogical_count: 14→15,PostTraining=15,BETA=10, ALPHA=5— arithmétique cohérente avec l'insertion d'une ligne. - Security scan : 1 match =
tokenizer.pad_token = tokenizer.eos_token— faux positif (assignation pad token, pas un credential). 0 vrai match (HF_TOKEN|API_KEY|BEARER|PASSWORD|SECRET|TOKEN=).
Structure conforme à la spec #12653, acceptances 2-3 explicitement gated sur c.558/c.559 — posture honnête. (Contrainte token : COMMENT only, auteur = jsboige.)
MD hierarchy drift -- 6a62105Cette PR augmente le compte de defauts de rendu markdown Corriger (ex. |
Golden-Set Execution (H.7 P3)✅ 8/8 notebooks passed (certified reproducible)
Pinned lockfile: |
|
[CONTEXTE REVIEW c.563] PR #12885 — MD hierarchy drift advisory : faux positif scanner (c.549-L1 ★ NEW pattern) ContexteLane Diagnostic : FAUX POSITIF SCANNER (c.549-L1 ★ NEW)La PR #12885 ne touche AUCUN des 3 notebooks signalés. Vérification firsthand Seuls fichiers modifiés :
Aucun des 3 notebooks signalés ( Tell = c.549-L1 ★ NEWLe scanner MD hierarchy drift ne filtre pas le diff par PR : il compare la baseline (burndown) au corpus complet, et tout delta — même sur des fichiers que la PR ne touche pas — est attribué à la PR. Pattern mesuré c.549 sur #12825 ( Note :
Note : le commentaire Action
Refs
Lane : myia-po-2026:CoursIA-2 |
MD hierarchy drift -- 5353f8cCette PR augmente le compte de defauts de rendu markdown Corriger (ex. |
MD hierarchy drift -- 0bffacfCette PR augmente le compte de defauts de rendu markdown Corriger (ex. |
Notebook PR Validation: PASS
Checks: H.1 (no errors), H.3 (execution_count), C.1 (no banned patterns) |
feat(post-training,#12653): PT-11c GRPO-RLVR cran 1.7B/2B — structure notebook
Grain: DEEP/training -- lane myia-po-2026:CoursIA-2 -- prev: LIGHT/refactor (cycles narrow sustained c.512+)
Résumé
Création structurelle du notebook
PT_11c_grpo_qwen17_rlvr.ipynbsur base PT-11b, avec adaptations Qwen3-1.7B/2B (cran au-dessus 0.8B). Cycle c.557 : structure ; c.558 : run live 4 seeds × 100 steps sur RTX 3080 Ti 16GB ; c.559 : verdict honnête + post-merge livraison.Livrables
MyIA.AI.Notebooks/GenAI/PostTraining/PT_11c_grpo_qwen17_rlvr.ipynb(33 cellules, 18 md + 15 code)Architecture
trl.GRPOTrainerextract_answer_sympy+math_verifier_rewardrewardspy.watch_trl(optionnel) +rlvr_reward_func#10603loss_fn='linear'informatif_torch.cuda.max_memory_allocated(0) / 1e9Acceptances #12653
linear+ delta > 0Verdict 4 classes (cf
pr-review-discipline.md §C) :linearp<0.05 ET delta > 0Règles respectées
raise NotImplementedError/assert False/1/0, stubspass/return Noneconformes (3 exos)Hors-scope ce cycle
Voir aussi
sensitivity_lean+knot_leannon-buildable — pas concerné ici (PT-11c est training Python)See #12653 (livraison partielle — issue reste ouverte pour c.558/c.559)