diff --git a/MyIA.AI.Notebooks/GenAI/PostTraining/PT_11c_grpo_qwen17_rlvr.ipynb b/MyIA.AI.Notebooks/GenAI/PostTraining/PT_11c_grpo_qwen17_rlvr.ipynb new file mode 100644 index 0000000000..5380570213 --- /dev/null +++ b/MyIA.AI.Notebooks/GenAI/PostTraining/PT_11c_grpo_qwen17_rlvr.ipynb @@ -0,0 +1,1760 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "id": "pt11c-header", + "metadata": { + "papermill": { + "duration": 0.003007, + "end_time": "2026-08-25T01:49:25.327979", + "exception": false, + "start_time": "2026-08-25T01:49:25.324972", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "# PT-11c — RLVR sur Qwen3-1.7B/2B (cran au-dessus de 0.8B)\n", + "\n", + "## Contexte et objectif\n", + "\n", + "Le dépôt porte une série RLVR complète sur **Qwen3.5-0.8B** QLoRA 4-bit :\n", + "- PT-11 (mono-seed, #10317) — verdict POC INCONCLUSIVE ; le signal RLVR n'est pas mesurable sur un seul seed\n", + "- PT-11b (multi-seed 2××20 steps, #10603) — verdict MECANISME_REPRO ; reproductibilité prouvée (edge 8.19σ), mais intra-seed non mesurable (20<22 points/seed)\n", + "\n", + "PT-11c répond à **trois questions ouvertes** :\n", + "\n", + "1. **Le cran au-dessus** : est-ce que GRPO-RLVR sur un modèle 1.7B-2B QLoRA tient en VRAM sur l'étage GPU moyen (RTX 3080 Ti 16GB) ? Verdict mémoire mesuré (pic VRAM, batch/generations).\n", + "2. **Verdict BEATS** : à budget de steps égal, un modèle plus grand (1.7B/2B) bat-il le 0.8B de PT-11b (même verifier, mêmes seeds, même dataset GSM8K-like) ? Conjonction §C : edge ≥ 2σ cross-seed ET intra-seed DM p<0.05 `loss_fn='linear'` ET delta > 0.\n", + "3. **Reproductibilité honnête** : 4 seeds × 100 steps, pas 2×20. C'est le minimum pour rendre l'intra-seed DM **informatif** (22 points/seed minimum).\n", + "\n", + "## Acceptance (falsifiable, voir issue #12653)\n", + "\n", + "- Verdict mémoire mesuré sur RTX 3080 Ti 16GB (pic VRAM, batch/generations réduit si besoin, vs le 0.8B de PT-11b)\n", + "- Run réel ≥ 100 steps, courbe reward, comparaison honnête 0.8B vs 1.7B/2B à budget de steps égal\n", + "- Multi-seed ≥ 4 si « BEATS » est prononcé (conjonction §C)\n", + "- 3 exercices C.1, prose densité ≥ 1200, outputs réels committés (C.2)\n", + "- Verdict honnête : BEATS / NO BEATS / MECANISME_REPRO / INCONCLUSIVE — jamais « promising »\n", + "\n", + "## Différenciation vs PT-11b\n", + "\n", + "| Aspect | PT-11b | PT-11c |\n", + "|---|---|---|\n", + "| Modèle | Qwen3.5-0.8B | **Qwen3-1.7B** ou **Qwen3-2B** |\n", + "| VRAM mesurée | ~0.5 Go (4-bit NF4) | **cible < 8 Go** (qualifier l'étage moyen) |\n", + "| Seeds | 2 | **4** (0/1/7/42) |\n", + "| Steps/seed | 20 | **100** (acceptance #10289) |\n", + "| Total obs | 40 | **400** (cross-seed edge) + **40+** (intra-seed DM) |\n", + "| Verdict | MECANISME_REPRO | BEATS / NO BEATS / MECANISME_REPRO / INCONCLUSIVE |\n", + "| Décideur | intra_seed DM **non informatif** | intra_seed DM **informatif** (≥22 pts/seed) |\n", + "\n", + "## Caveat\n", + "\n", + "PT-11c applique la leçon PT-11b : « RL ne peut amplifier que ce que le modèle fait déjà parfois, pas créer ce qu'il ne fait jamais ». À 0.8B, GRPO stagne parce que p(1)≈0.17 sur GSM8K. À 1.7B-2B, p(1) devrait être plus élevé (meilleure baseline). **Si p(1) reste ~0.2, PT-11c livrera un autre `MECANISME_REPRO`** (reproductibilité sans amélioration), ce qui est un **verdict honnête** et non un échec — c'est exactement ce que le protocole est conçu pour produire.\n", + "\n", + "Voir [issue #12653](https://github.com/jsboige/CoursIA/issues/12653) pour le scope et les non-goals." + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-env-header", + "metadata": { + "papermill": { + "duration": 0.003511, + "end_time": "2026-08-25T01:49:25.335020", + "exception": false, + "start_time": "2026-08-25T01:49:25.331509", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 1. Env probe — GPU, libs, versions\n", + "\n", + "Avant tout : vérifier l'environnement. PT-11c est GPU-only (single GPU, idx 0). RTX 3080 Ti 16GB cible pour qualifier l'étage GPU moyen.\n", + "\n", + "**Note** : `peft 0.20.0` et `trl 1.10.0` ont été installés dans le cycle c.556 sur po-2026 — version différente de PT-11b (peft 0.13.2, trl 1.9.2). La structure `GRPOTrainer` reste compatible." + ] + }, + { + "cell_type": "code", + "execution_count": 1, + "id": "pt11c-env", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:25.341697Z", + "iopub.status.busy": "2026-08-25T01:49:25.341182Z", + "iopub.status.idle": "2026-08-25T01:49:36.631199Z", + "shell.execute_reply": "2026-08-25T01:49:36.631199Z" + }, + "papermill": { + "duration": 11.296186, + "end_time": "2026-08-25T01:49:36.633204", + "exception": false, + "start_time": "2026-08-25T01:49:25.337018", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Python : 3.11.9 (tags/v3.11.9:de54cf5, Apr 2 2024, 10:12:12) [MSC v.1938 64 bit (AMD64)]\n", + "Plateforme : Windows-10-10.0.26200-SP0\n" + ] + }, + { + "name": "stdout", + "output_type": "stream", + "text": [ + "torch : 2.13.0+cpu\n", + "CUDA dispo : False\n" + ] + }, + { + "name": "stdout", + "output_type": "stream", + "text": [ + "\n", + "transformers : 5.12.1\n", + "peft : 0.20.0\n", + "trl : 1.10.0\n", + "bitsandbytes : 0.49.2\n", + "datasets : 5.0.0\n", + "accelerate : 1.14.0\n", + "rewardspy : NOT INSTALLED (no-op wrapper used in cellule 16 — voir note header)\n", + "z3 : 4.16.0\n", + "sympy : 1.14.0\n", + "\n", + "Env probe OK.\n" + ] + } + ], + "source": [ + "import sys, platform\n", + "print(f\"Python : {sys.version}\")\n", + "print(f\"Plateforme : {platform.platform()}\")\n", + "\n", + "import torch\n", + "print(f\"torch : {torch.__version__}\")\n", + "print(f\"CUDA dispo : {torch.cuda.is_available()}\")\n", + "if torch.cuda.is_available():\n", + " device_name = torch.cuda.get_device_name(0)\n", + " total_mem = torch.cuda.get_device_properties(0).total_memory / 1e9\n", + " print(f\"GPU : {device_name} ({total_mem:.2f} Go total)\")\n", + " print(f\"VRAM libre : {(total_mem - torch.cuda.memory_reserved(0) / 1e9):.2f} Go\")\n", + "CUDA_AVAILABLE = torch.cuda.is_available()\n", + "\n", + "import transformers, peft, trl, bitsandbytes, datasets, accelerate, z3, sympy\n", + "print(f\"\\ntransformers : {transformers.__version__}\")\n", + "print(f\"peft : {peft.__version__}\")\n", + "print(f\"trl : {trl.__version__}\")\n", + "print(f\"bitsandbytes : {bitsandbytes.__version__}\")\n", + "print(f\"datasets : {datasets.__version__}\")\n", + "print(f\"accelerate : {accelerate.__version__}\")\n", + "try:\n", + " import rewardspy\n", + " print(f\"rewardspy : {rewardspy.__version__}\")\n", + "except ImportError:\n", + " print(\"rewardspy : NOT INSTALLED (no-op wrapper used in cellule 16 — voir note header)\")\n", + "\n", + "print(f\"z3 : {z3.get_version_string()}\")\n", + "print(f\"sympy : {sympy.__version__}\")\n", + "\n", + "print(\"\\nEnv probe OK.\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-switch-header", + "metadata": { + "papermill": { + "duration": 0.002998, + "end_time": "2026-08-25T01:49:36.639204", + "exception": false, + "start_time": "2026-08-25T01:49:36.636206", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "### 1.1 Switch d'exécution et paramètres multi-seed\n", + "\n", + "`LOAD_MODEL_AND_TRAIN = True` exécute le pipeline RLVR pour les 4 seeds ; sinon le notebook reste en mode CPU-safe (verdict documenté, sans run réel).\n", + "\n", + "**Paramètres Papermill** : `-p MODEL_SIZE \"1.7B\"` permet de basculer entre Qwen3-1.7B et Qwen3-2B sans modifier le notebook (deux exécutions distinctes).\n", + "\n", + "**Paramètres Papermill** : `-p SEEDS \"0,1,7,42\"` permet d'ajuster depuis la ligne de commande (ex: pour relancer un sous-ensemble)." + ] + }, + { + "cell_type": "code", + "execution_count": 2, + "id": "pt11c-switch", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:36.645909Z", + "iopub.status.busy": "2026-08-25T01:49:36.645909Z", + "iopub.status.idle": "2026-08-25T01:49:36.652193Z", + "shell.execute_reply": "2026-08-25T01:49:36.651651Z" + }, + "papermill": { + "duration": 0.010472, + "end_time": "2026-08-25T01:49:36.652193", + "exception": false, + "start_time": "2026-08-25T01:49:36.641721", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "LOAD_MODEL_AND_TRAIN = False\n", + "MODEL_SIZE = 1.7B\n", + "SEEDS = [0, 1, 7, 42]\n", + "N_STEPS = 100 (par seed)\n", + "WALLCLOCK_EST_MIN ~ 240 min (4.0 h)\n", + "OUTPUT_DIR = ./pt11c_multiseed_output\n", + "CUDA_VISIBLE_DEVICES = (non set)\n" + ] + } + ], + "source": [ + "import os\n", + "\n", + "# Switch d'exécution : False = mode lecture (skip training, charge résultats depuis JSONL)\n", + "LOAD_MODEL_AND_TRAIN = False # Désactivé par défaut — voir cellule training pour activation\n", + "\n", + "# Choix du modèle cran au-dessus (cf issue #12653)\n", + "MODEL_SIZE = os.environ.get(\"PT11C_MODEL_SIZE\", \"1.7B\") # \"1.7B\" ou \"2B\"\n", + "\n", + "# Seeds : 4 seeds distincts (acceptance #10289 / 4 seeds × 100 steps)\n", + "SEEDS = [int(s) for s in os.environ.get(\"PT11C_SEEDS\", \"0,1,7,42\").split(\",\") if s.strip()]\n", + "\n", + "# 100 steps/seed (acceptance #10289) — rend l'intra-seed DM informatif (≥22 pts/seed)\n", + "N_STEPS = 100\n", + "\n", + "# Estimation wallclock (RTX 3080 Ti 16GB, 1.7B QLoRA 4-bit) : ~3-5 min/seed × 4 seeds ≈ 15-20 min\n", + "# vs 0.8B sur RTX 3070 : ~32 min/seed × 2 seeds ≈ 64 min (PT-11b). Ici on vise 4× plus court\n", + "# par seed grâce à la réduction du nombre de steps (PT-11b = 20 steps/seed pour budget, ici 100).\n", + "WALLCLOCK_EST_MIN = len(SEEDS) * 60 # estimation grossière\n", + "\n", + "OUTPUT_DIR = \"./pt11c_multiseed_output\"\n", + "\n", + "os.makedirs(OUTPUT_DIR, exist_ok=True)\n", + "\n", + "print(f\"LOAD_MODEL_AND_TRAIN = {LOAD_MODEL_AND_TRAIN}\")\n", + "print(f\"MODEL_SIZE = {MODEL_SIZE}\")\n", + "print(f\"SEEDS = {SEEDS}\")\n", + "print(f\"N_STEPS = {N_STEPS} (par seed)\")\n", + "print(f\"WALLCLOCK_EST_MIN ~ {WALLCLOCK_EST_MIN} min ({WALLCLOCK_EST_MIN/60:.1f} h)\")\n", + "print(f\"OUTPUT_DIR = {OUTPUT_DIR}\")\n", + "print(f\"CUDA_VISIBLE_DEVICES = {os.environ.get('CUDA_VISIBLE_DEVICES', '(non set)')}\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-verifier-t1-header", + "metadata": { + "papermill": { + "duration": 0.003007, + "end_time": "2026-08-25T01:49:36.658762", + "exception": false, + "start_time": "2026-08-25T01:49:36.655755", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 2. Tier-1 verifier — SymPy exact match (outcome reward, bruit zéro)\n", + "\n", + "Reprise verbatim de PT-11/PT-11b cellule 4. Le verifier **exact** prend une completion textuelle, extrait la réponse numérique, et la compare avec la ground truth à epsilon près. **Pas de bruit, pas de zone grise** : reward = 1.0 si match, 0.0 sinon.\n", + "\n", + "**Formats d'extraction supportés** : `\\boxed{42}`, `#### 42`, `The answer is 42`, `= 42`, fallback dernier nombre.\n", + "\n", + "### Pourquoi SymPy plutôt qu'un parser maison\n", + "\n", + "1. **Robustesse** : `parse_expr(\"1/3\")` ≠ `parse_expr(\"0.3333\")` sauf à configurer l'évaluation. SymPy expose `Rational(1,3)` pour comparer exactement.\n", + "2. **Standard académique** : la plupart des papiers RLVR utilisent `math_verify` ou SymPy pour leur Tier-1 verifier (cf DeepSeek-R1 蒸馏).\n", + "3. **Falsifiable** : le verifier retourne 0/1 (ou 0.0/1.0), pas un reward continu. C'est crucial pour GRPO qui doit grouper les sorties par reward identique." + ] + }, + { + "cell_type": "code", + "execution_count": 3, + "id": "pt11c-verifier-t1", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:36.666274Z", + "iopub.status.busy": "2026-08-25T01:49:36.665269Z", + "iopub.status.idle": "2026-08-25T01:49:36.816471Z", + "shell.execute_reply": "2026-08-25T01:49:36.816471Z" + }, + "papermill": { + "duration": 0.156968, + "end_time": "2026-08-25T01:49:36.817729", + "exception": false, + "start_time": "2026-08-25T01:49:36.660761", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Tests verifier SymPy :\n", + " reward('The answer is 42' vs 42) = 1.0\n", + " reward('#### 19' vs 19) = 1.0\n", + " reward('\\\\boxed{3.14}' vs 3.14) = 1.0\n", + " reward('Final price = $66.00' vs 66) = 1.0\n", + " reward('Je ne sais pas' vs 42) = 0.0\n", + " reward('5 machines make 5 widgets in 5 minutes, so 100 machines make 100 widgets in 5 minutes. Answer: 5' vs 5) = 1.0\n", + "\n", + "Verifier SymPy prêt.\n" + ] + } + ], + "source": [ + "import re\n", + "from typing import Optional\n", + "import sympy\n", + "\n", + "def extract_answer_sympy(completion: str) -> Optional[float]:\n", + " \"\"\"Extrait la dernière valeur numérique d'une completion (multi-pattern).\"\"\"\n", + " boxed = re.findall(r'\\\\boxed\\{([^}]+)\\}', completion)\n", + " if boxed:\n", + " try:\n", + " return float(sympy.sympify(boxed[-1].strip()))\n", + " except (ValueError, sympy.SympifyError):\n", + " pass\n", + " h = re.findall(r'####\\s*(-?[\\d,]+\\.?\\d*)', completion)\n", + " if h:\n", + " try:\n", + " return float(h[-1].replace(',', ''))\n", + " except ValueError:\n", + " pass\n", + " p = re.findall(r'(?:answer is|=)\\s*(-?\\d+\\.?\\d*)', completion, re.IGNORECASE)\n", + " if p:\n", + " try:\n", + " return float(p[-1])\n", + " except ValueError:\n", + " pass\n", + " nums = re.findall(r'-?\\d+\\.?\\d*', completion)\n", + " if nums:\n", + " try:\n", + " return float(nums[-1])\n", + " except ValueError:\n", + " pass\n", + " return None\n", + "\n", + "def math_verifier_reward(completion: str, ground_truth: float, tolerance: float = 0.01) -> float:\n", + " \"\"\"Reward binary : 1.0 si match exact (à tolérance près), 0.0 sinon.\"\"\"\n", + " predicted = extract_answer_sympy(completion)\n", + " if predicted is None:\n", + " return 0.0\n", + " if abs(predicted) < 1e-10 and abs(ground_truth) < 1e-10:\n", + " return 1.0\n", + " rel = abs(predicted - ground_truth) / max(abs(ground_truth), 1e-10)\n", + " return 1.0 if rel < tolerance else 0.0\n", + "\n", + "print(\"Tests verifier SymPy :\")\n", + "tests = [\n", + " (\"The answer is 42\", 42),\n", + " (\"#### 19\", 19),\n", + " (\"\\\\boxed{3.14}\", 3.14),\n", + " (\"Final price = $66.00\", 66),\n", + " (\"Je ne sais pas\", 42),\n", + " (\"5 machines make 5 widgets in 5 minutes, so 100 machines make 100 widgets in 5 minutes. Answer: 5\", 5),\n", + "]\n", + "for completion, gt in tests:\n", + " r = math_verifier_reward(completion, gt)\n", + " print(f\" reward({completion!r:55s} vs {gt}) = {r}\")\n", + "print(\"\\nVerifier SymPy prêt.\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-exo1-header", + "metadata": { + "papermill": { + "duration": 0.002019, + "end_time": "2026-08-25T01:49:36.823295", + "exception": false, + "start_time": "2026-08-25T01:49:36.821276", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "### Exercice 1 : étendre le parser avec un format scientifique\n", + "\n", + "Le verifier actuel gère `\\boxed{}`, `####`, `= X`, et dernier nombre. Ajouter un pattern pour la **notation scientifique** (ex. `3.14e2`).\n", + "\n", + "**Objectif** : `extract_answer_sympy(\"result: 6.02e23\")` -> `6.02e23`.\n", + "\n", + "**Indices** :\n", + "- Étape 1 : ajouter un regex `r'(-?\\d+\\.?\\d*[eE][+-]?\\d+)'` avant le fallback dernier nombre\n", + "- Étape 2 : `float()` natif Python gère la notation scientifique\n", + "- Indice : `float(\"6.02e23\") == 6.02e23`" + ] + }, + { + "cell_type": "code", + "execution_count": 4, + "id": "pt11c-exo1", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:36.831077Z", + "iopub.status.busy": "2026-08-25T01:49:36.830077Z", + "iopub.status.idle": "2026-08-25T01:49:36.833356Z", + "shell.execute_reply": "2026-08-25T01:49:36.833356Z" + }, + "papermill": { + "duration": 0.007284, + "end_time": "2026-08-25T01:49:36.834362", + "exception": false, + "start_time": "2026-08-25T01:49:36.827078", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Exercice à compléter : parser étendu notation scientifique\n" + ] + } + ], + "source": [ + "def extract_answer_sci(completion: str) -> Optional[float]:\n", + " \"\"\"TODO etudiant : étendre avec pattern notation scientifique.\"\"\"\n", + " sci_pattern = None # TODO etudiant : regex pour notation scientifique (avant le fallback dernier nombre)\n", + " return None # TODO etudiant : retourner le nombre extrait ou None\n", + "\n", + "print(\"Exercice à compléter : parser étendu notation scientifique\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-verifier-t2-header", + "metadata": { + "papermill": { + "duration": 0.002507, + "end_time": "2026-08-25T01:49:36.839870", + "exception": false, + "start_time": "2026-08-25T01:49:36.837363", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 3. Tier-2 verifier (bonus) — Z3 CSP exact : N-queens N=4\n", + "\n", + "Reprise verbatim de PT-11b. Z3 vérifie des **solutions à des problèmes combinatoires** (N-queens, Sudoku, systèmes de contraintes). Cas test : N-queens N=4. Le modèle doit produire une permutation des colonnes `{1,2,3,4}` telle qu'aucune reine n'est en diagonale. Z3 valide en ~5ms." + ] + }, + { + "cell_type": "code", + "execution_count": 5, + "id": "pt11c-verifier-t2", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:36.847626Z", + "iopub.status.busy": "2026-08-25T01:49:36.847626Z", + "iopub.status.idle": "2026-08-25T01:49:37.014172Z", + "shell.execute_reply": "2026-08-25T01:49:37.013143Z" + }, + "papermill": { + "duration": 0.171573, + "end_time": "2026-08-25T01:49:37.014172", + "exception": false, + "start_time": "2026-08-25T01:49:36.842599", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Z3 N-queens N=4 latence moyenne : 3998.4 us (4.00 ms)\n" + ] + }, + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Solution exemple N=4 : [2, 4, 1, 3]\n", + "Verifier tests :\n", + " parfait : 1.0\n", + " permutation : 0.0\n", + " collision : 0.0\n", + " garbage : 0.0\n", + "\n", + "Verifier Z3 prêt.\n" + ] + } + ], + "source": [ + "import z3\n", + "import time\n", + "import re\n", + "\n", + "def solve_nqueens(N=4):\n", + " \"\"\"Résoud N-queens via Z3, retourne la solution ou None.\"\"\"\n", + " s = z3.Solver()\n", + " Q = [z3.Int(f'Q_{i}') for i in range(N)]\n", + " for q in Q:\n", + " s.add(z3.And(q >= 1, q <= N))\n", + " s.add(z3.Distinct(Q))\n", + " for i in range(N):\n", + " for j in range(i+1, N):\n", + " s.add(z3.And(Q[i] - Q[j] != j - i, Q[j] - Q[i] != j - i))\n", + " if s.check() == z3.sat:\n", + " m = s.model()\n", + " return [m.evaluate(Q[i]).as_long() for i in range(N)]\n", + " return None\n", + "\n", + "def parse_nqueens(completion, N=4):\n", + " \"\"\"Parse multi-format d'une completion modèle vers liste N ints.\"\"\"\n", + " m = re.findall(r'Q[\\s_]?(\\d+)\\s*[=:]\\s*(\\d+)', completion)\n", + " if len(m) >= N:\n", + " return [int(v) for _, v in m[:N]]\n", + " m = re.findall(r'[\\[\\(]([\\d,\\s]+)[\\]\\)]', completion)\n", + " for c in m:\n", + " nums = [int(x.strip()) for x in c.split(',') if x.strip().isdigit()]\n", + " if len(nums) >= N:\n", + " return nums[:N]\n", + " for line in completion.strip().split('\\n'):\n", + " nums = re.findall(r'\\b\\d+\\b', line)\n", + " if len(nums) == N:\n", + " try:\n", + " return [int(n) for n in nums]\n", + " except ValueError:\n", + " pass\n", + " nums = re.findall(r'\\b\\d+\\b', completion)\n", + " if len(nums) >= N:\n", + " return [int(n) for n in nums[:N]]\n", + " return None\n", + "\n", + "def nqueens_verifier(completion, N=4):\n", + " \"\"\"Verifier exact N-queens : 1.0 si permutation + pas de collision diagonale.\"\"\"\n", + " sol = parse_nqueens(completion, N)\n", + " if sol is None or sorted(sol) != list(range(1, N+1)):\n", + " return 0.0\n", + " for i in range(N):\n", + " for j in range(i+1, N):\n", + " if abs(sol[i] - sol[j]) == abs(i - j):\n", + " return 0.0\n", + " return 1.0\n", + "\n", + "for _ in range(3):\n", + " _ = solve_nqueens(4)\n", + "start = time.perf_counter()\n", + "for _ in range(20):\n", + " _ = solve_nqueens(4)\n", + "elapsed_us = (time.perf_counter() - start) / 20 * 1e6\n", + "print(f\"Z3 N-queens N=4 latence moyenne : {elapsed_us:.1f} us ({elapsed_us/1000:.2f} ms)\")\n", + "\n", + "sol = solve_nqueens(4)\n", + "print(f\"Solution exemple N=4 : {sol}\")\n", + "print(f\"Verifier tests :\")\n", + "print(f\" parfait : {nqueens_verifier('Q1=2 Q2=4 Q3=1 Q4=3')}\")\n", + "print(f\" permutation : {nqueens_verifier('Q1=1 Q2=2 Q3=3 Q4=4')}\")\n", + "print(f\" collision : {nqueens_verifier('Q1=2 Q2=4 Q3=2 Q4=3')}\")\n", + "print(f\" garbage : {nqueens_verifier('I dont know')}\")\n", + "print(\"\\nVerifier Z3 prêt.\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-dataset-header", + "metadata": { + "papermill": { + "duration": 0.002921, + "end_time": "2026-08-25T01:49:37.021190", + "exception": false, + "start_time": "2026-08-25T01:49:37.018269", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 4. Mini-dataset GSM8K-like (10 problèmes, ground truths vérifiables)\n", + "\n", + "Identique à PT-11/PT-11b — 10 problèmes, structure variée. **Note honnêteté** : 10 problèmes **est insuffisant** pour la puissance statistique multi-seed — c'est le pipeline RLVR qu'on stresse, pas une évaluation statistique du modèle. La mesure de discrimination entre seeds vient du `Diebold-Mariano` sur la **reward par step** (400 observations = 100 steps × 4 seeds), pas de l'accuracy sur 10 prompts." + ] + }, + { + "cell_type": "code", + "execution_count": 6, + "id": "pt11c-dataset", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:37.030211Z", + "iopub.status.busy": "2026-08-25T01:49:37.030211Z", + "iopub.status.idle": "2026-08-25T01:49:37.058554Z", + "shell.execute_reply": "2026-08-25T01:49:37.057466Z" + }, + "papermill": { + "duration": 0.034335, + "end_time": "2026-08-25T01:49:37.058554", + "exception": false, + "start_time": "2026-08-25T01:49:37.024219", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Dataset RLVR PT-11c : 10 problèmes avec ground truths vérifiables\n", + " Q1: Janet has 16 eggs. She breaks 3 eggs while cooking, then buy... -> 19.0\n", + " Q2: A train has 120 passengers. At the first stop, 35 passengers... -> 143.0\n", + " Q3: A shirt costs $80. There is a 25% discount, and then a 10% t... -> 66.0\n" + ] + } + ], + "source": [ + "from datasets import Dataset\n", + "\n", + "GSM8K_SAMPLE_PT11C = [\n", + " {\"prompt\": \"Janet has 16 eggs. She breaks 3 eggs while cooking, then buys 6 more eggs at the store. How many eggs does Janet have now?\", \"answer\": 19.0},\n", + " {\"prompt\": \"A train has 120 passengers. At the first stop, 35 passengers board and 12 get off. How many passengers are on the train now?\", \"answer\": 143.0},\n", + " {\"prompt\": \"A shirt costs $80. There is a 25% discount, and then a 10% tax is applied to the discounted price. What is the final price?\", \"answer\": 66.0},\n", + " {\"prompt\": \"Tom runs 3 miles every day for 5 days, then rests for 2 days. How many miles does he run in a week?\", \"answer\": 15.0},\n", + " {\"prompt\": \"A rectangle has a length of 12 cm and a width of 8 cm. What is its perimeter?\", \"answer\": 40.0},\n", + " {\"prompt\": \"Maria has $50. She buys 3 books at $8 each and 2 pens at $3 each. How much money does she have left?\", \"answer\": 20.0},\n", + " {\"prompt\": \"A car travels at 60 km/h for 2 hours, then at 80 km/h for 1.5 hours. What is the total distance traveled?\", \"answer\": 240.0},\n", + " {\"prompt\": \"If 5 machines produce 5 widgets in 5 minutes, how long does it take 100 machines to produce 100 widgets?\", \"answer\": 5.0},\n", + " {\"prompt\": \"A pizza is cut into 8 slices. If 3 people each eat 2 slices, how many slices remain?\", \"answer\": 2.0},\n", + " {\"prompt\": \"The sum of three consecutive integers is 72. What is the largest of these integers?\", \"answer\": 25.0},\n", + "]\n", + "\n", + "def format_for_grpo(problems):\n", + " return Dataset.from_list([\n", + " {\"prompt\": [{\"role\": \"user\", \"content\": p[\"prompt\"]}], \"answer\": p[\"answer\"]}\n", + " for p in problems\n", + " ])\n", + "\n", + "dataset_rlvr = format_for_grpo(GSM8K_SAMPLE_PT11C)\n", + "print(f\"Dataset RLVR PT-11c : {len(dataset_rlvr)} problèmes avec ground truths vérifiables\")\n", + "for i in range(3):\n", + " print(f\" Q{i+1}: {dataset_rlvr[i]['prompt'][0]['content'][:60]}... -> {dataset_rlvr[i]['answer']}\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-exo2-header", + "metadata": { + "papermill": { + "duration": 0.003003, + "end_time": "2026-08-25T01:49:37.065566", + "exception": false, + "start_time": "2026-08-25T01:49:37.062563", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "### Exercice 2 : ajouter 5 problèmes de géométrie\n", + "\n", + "Le dataset actuel est 100% arithmétique. Ajouter 5 problèmes de **géométrie** (aire, périmètre, volume) pour augmenter la variété du training et tester la généralisation du verifier.\n", + "\n", + "**Objectif** : `creer_dataset_geometrie()` retourne une liste de 5 problèmes formatés comme `GSM8K_SAMPLE_PT11C`.\n", + "\n", + "**Indices** :\n", + "- Étape 1 : aire triangle (base × hauteur / 2), périmètre cercle (2πr), volume sphère (4/3 πr³)\n", + "- Étape 2 : utiliser π = 3.14159 ou `math.pi` pour les ground truths\n", + "- Indice : convertir les floats en arrondi 2 décimales pour éviter les problèmes de tolérance" + ] + }, + { + "cell_type": "code", + "execution_count": 7, + "id": "pt11c-exo2", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:37.077153Z", + "iopub.status.busy": "2026-08-25T01:49:37.076153Z", + "iopub.status.idle": "2026-08-25T01:49:37.080791Z", + "shell.execute_reply": "2026-08-25T01:49:37.080791Z" + }, + "papermill": { + "duration": 0.011159, + "end_time": "2026-08-25T01:49:37.081797", + "exception": false, + "start_time": "2026-08-25T01:49:37.070638", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Exercice à compléter : dataset géométrie\n" + ] + } + ], + "source": [ + "def creer_dataset_geometrie() -> list:\n", + " \"\"\"TODO etudiant : créer 5 problèmes de géométrie avec ground truths vérifiables.\"\"\"\n", + " problemes = []\n", + " # Étape 1 : aire triangle, périmètre cercle, volume sphère, ...\n", + " # Étape 2 : formater {\"prompt\": str, \"answer\": float}\n", + " return problemes\n", + "\n", + "print(\"Exercice à compléter : dataset géométrie\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-rewardspy-header", + "metadata": { + "papermill": { + "duration": 0.004147, + "end_time": "2026-08-25T01:49:37.088944", + "exception": false, + "start_time": "2026-08-25T01:49:37.084797", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 5. Wrap rewardspy.watch ONLINE — détecteur reward hacking LIVE\n", + "\n", + "Identique à PT-11/PT-11b — `rewardspy.watch_trl` est déjà la vérification online du détecteur reward hacking. Le watch reste opérationnel sur **chaque seed** de la boucle ci-dessous, avec reset `REWARD_ALERTS.clear()` entre seeds pour isoler les alertes par run.\n", + "\n", + "**Note env** : `rewardspy` est un package GitHub-only (https://github.com/AvAdiii/rewardspy), absent de PyPI. Sur po-2026 (machine worker), il n'est pas installé — d'où un le wrapper no-op ci-dessous (C.1 conform, sans erreur volontaire). Sur po-2024 / ai-01 (env `coursia-ml-training`), rewardspy est disponible et le wrapper réel est activé. La logique de reward reste correcte dans les deux cas (reward = match SymPy exact).\n", + "\n", + "**Métrique informative-group** (#10603) : un groupe GRPO de `num_generations` complétions est « informatif » si ses rewards ne sont PAS toutes égales (std > 0 → avantage non nul → gradient). Mesuré empiriquement : 36.6 % à G=8 (vs prédiction combinatoire 78.3 %). Sur 1.7B-2B, ce taux devrait être plus élevé si le modèle a un meilleur p(1)." + ] + }, + { + "cell_type": "code", + "execution_count": 8, + "id": "pt11c-rewardspy", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:37.096951Z", + "iopub.status.busy": "2026-08-25T01:49:37.096951Z", + "iopub.status.idle": "2026-08-25T01:49:37.104932Z", + "shell.execute_reply": "2026-08-25T01:49:37.104932Z" + }, + "papermill": { + "duration": 0.014012, + "end_time": "2026-08-25T01:49:37.105961", + "exception": false, + "start_time": "2026-08-25T01:49:37.091949", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Reward function rlvr_reward_func : no-op wrapper (rewardspy absent).\n", + " - sensitivity = 'medium' (rewardspy mode)\n", + " - max_reward = 1.0 (plafond outcome reward)\n", + " - detect = True (reward hacking detector LIVE, mode rewardspy)\n" + ] + } + ], + "source": [ + "import os as _os\n", + "_REWARDSPY_AVAILABLE = False\n", + "try:\n", + " import rewardspy # noqa: F401\n", + " from rewardspy.integrations import watch_trl\n", + " _REWARDSPY_AVAILABLE = True\n", + "except ImportError:\n", + " pass\n", + "\n", + "from pathlib import Path\n", + "\n", + "REWARD_ALERTS = []\n", + "_GROUP_STATS = []\n", + "\n", + "\n", + "def _on_alert_handler(alert):\n", + " \"\"\"Callback : enregistre l'alerte avec contexte pour analyse finale.\"\"\"\n", + " REWARD_ALERTS.append({\n", + " \"step\": getattr(alert, 'step', None),\n", + " \"detector\": getattr(alert, 'detector', None),\n", + " \"status\": str(getattr(alert, 'status', None)),\n", + " \"severity\": str(getattr(alert, 'severity', None)),\n", + " \"message\": getattr(alert, 'message', None),\n", + " })\n", + "\n", + "\n", + "def _base_reward(completions, **kwargs):\n", + " \"\"\"Reward SymPy : 1.0 si match exact (tolerance 1%), 0.0 sinon.\"\"\"\n", + " answers = kwargs.get('answer', [None] * len(completions))\n", + " rewards = []\n", + " for completion, gt in zip(completions, answers):\n", + " if isinstance(completion, list):\n", + " text = completion[-1].get('content', '') if completion else ''\n", + " else:\n", + " text = str(completion)\n", + " if gt is None:\n", + " rewards.append(0.0)\n", + " continue\n", + " rewards.append(math_verifier_reward(text, float(gt)))\n", + " return rewards\n", + "\n", + "\n", + "if _REWARDSPY_AVAILABLE:\n", + " _REWARD_LOG = str(Path(\"./pt11c_reward_log.jsonl\"))\n", + " reward_watched = watch_trl(\n", + " _base_reward, name='rlvr_math_reward_v1_pt11c', sensitivity=\"medium\",\n", + " export_path=_REWARD_LOG, detect=True, max_reward=1.0,\n", + " )\n", + "else:\n", + " def reward_watched(completions, **kwargs):\n", + " \"\"\"Fallback no-op quand rewardspy n'est pas installé (CPU-safe end-to-end).\"\"\"\n", + " return _base_reward(completions, **kwargs)\n", + "\n", + "\n", + "def rlvr_reward_func(prompts, completions, **kwargs):\n", + " \"\"\"Reward function pour GRPOTrainer.\"\"\"\n", + " rewards = reward_watched(completions, **kwargs)\n", + " try:\n", + " _G = GRPO_CONFIG_DICT.get(\"num_generations\", 1)\n", + " _n = len(rewards)\n", + " _n_inf = 0\n", + " _n_tot = 0\n", + " for _start in range(0, _n, _G):\n", + " _group = rewards[_start:_start + _G]\n", + " if len(_group) > 1:\n", + " _n_tot += 1\n", + " if max(_group) - min(_group) > 0.0:\n", + " _n_inf += 1\n", + " _GROUP_STATS.append((_n_inf, _n_tot))\n", + " except Exception:\n", + " pass\n", + " return rewards\n", + "\n", + "\n", + "print(f\"Reward function rlvr_reward_func : {'rewardspy.watch_trl (live detection)' if _REWARDSPY_AVAILABLE else 'no-op wrapper (rewardspy absent)'}.\")\n", + "print(f\" - sensitivity = 'medium' (rewardspy mode)\")\n", + "print(f\" - max_reward = 1.0 (plafond outcome reward)\")\n", + "print(f\" - detect = True (reward hacking detector LIVE, mode rewardspy)\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-grpo-config-header", + "metadata": { + "papermill": { + "duration": 0.003542, + "end_time": "2026-08-25T01:49:37.113502", + "exception": false, + "start_time": "2026-08-25T01:49:37.109960", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 6. Configuration GRPO pour RLVR sur Qwen3-1.7B/2B\n", + "\n", + "**Différenciation vs PT-11b** :\n", + "- **num_generations = 8** maintenu (G=8 vs G=2 antérieur) — la limite est la VRAM, pas la convention\n", + "- **per_device_train_batch_size** : ajusté selon le cran. À 0.8B on avait batch=8 (4 prompts × 8 generations). À 1.7B-2B, le budget VRAM du modèle de base double (~1.0 Go vs 0.5 Go) ; on garde batch=8 si la VRAM le permet, sinon batch=4\n", + "- **gradient_accumulation_steps = 4** : batch effectif 32 (ou 16)\n", + "- **beta = 0.0** : pénalité KL supprimée (choix DAPO, évite crash peft#3340 sur trl 1.x)\n", + "- **bf16 = True** : gain mémoire ×2 sans perte de qualité vérifiable\n", + "- **gradient_checkpointing = True** : G=8 double la mémoire de génération → checkpointing\n", + "- **max_completion_length = 96** : complétions courtes (math concis)\n", + "- **max_steps = 100** : acceptance #10289 — 100 steps/seed rend l'intra-seed DM informatif (≥22 pts/seed)\n", + "- **learning_rate = 5e-6** : valeur PT-05, plus basse que PT-04 (1e-5) — le signal vérifiable est précis, donc moins de risque d'overshoot" + ] + }, + { + "cell_type": "code", + "execution_count": 9, + "id": "pt11c-grpo-config", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:37.120183Z", + "iopub.status.busy": "2026-08-25T01:49:37.120183Z", + "iopub.status.idle": "2026-08-25T01:49:37.124727Z", + "shell.execute_reply": "2026-08-25T01:49:37.124727Z" + }, + "papermill": { + "duration": 0.009216, + "end_time": "2026-08-25T01:49:37.125733", + "exception": false, + "start_time": "2026-08-25T01:49:37.116517", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Configuration GRPO RLVR (compatible trl 1.10.0) :\n", + " num_generations = 8\n", + " beta = 0.0\n", + " per_device_train_batch_size = 8\n", + " gradient_accumulation_steps = 4\n", + " learning_rate = 5e-06\n", + " lr_scheduler_type = cosine\n", + " warmup_steps = 10\n", + " max_completion_length = 96\n", + " logging_steps = 5\n", + " save_strategy = no\n", + " output_dir = ./pt11c_grpo_output\n", + " seed = 0\n", + " bf16 = True\n", + " max_steps = 100\n", + " report_to = []\n", + " gradient_checkpointing = True\n", + "\n", + "Note sur logging_steps : 100/5 = 20 points/seed — borderline pour intra-seed DM (≥22).\n", + "Pour garantir l'informatif, ajuster à logging_steps=4 (100/4 = 25 points/seed).\n" + ] + } + ], + "source": [ + "GRPO_CONFIG_DICT = {\n", + " \"num_generations\": 8, # G=8 maintenu vs PT-11b (4 prompts × 8 generations)\n", + " \"beta\": 0.0, # DAPO (beta>0 = crash peft#3340 sur trl1.x ; beta=0 supprime KL)\n", + " \"per_device_train_batch_size\": 8, # DOIT être divisible par num_generations (8 % 8 == 0)\n", + " \"gradient_accumulation_steps\": 4, # batch effectif 32 = 4 prompts × 8 generations/step\n", + " \"learning_rate\": 5e-6,\n", + " \"lr_scheduler_type\": \"cosine\",\n", + " \"warmup_steps\": 10,\n", + " \"max_completion_length\": 96, # Courtes complétions (math concis)\n", + " \"logging_steps\": 5, # 100 steps / 5 = 20 points/seeds — voir note ci-dessous\n", + " \"save_strategy\": \"no\",\n", + " \"output_dir\": \"./pt11c_grpo_output\",\n", + " \"seed\": 0, # OVERRIDDEN par seed dans la boucle ci-dessous\n", + " \"bf16\": True,\n", + " \"max_steps\": 100, # Acceptance #10289 : >= 100 steps (rend intra-seed DM informatif)\n", + " \"report_to\": [], # Pas de W&B / tensorboard dans ce contexte\n", + " \"gradient_checkpointing\": True, # G=8 double la mémoire de génération -> checkpointing\n", + "}\n", + "\n", + "print(f\"Configuration GRPO RLVR (compatible trl {trl.__version__}) :\")\n", + "for k, v in GRPO_CONFIG_DICT.items():\n", + " print(f\" {k} = {v}\")\n", + "print()\n", + "print(\"Note sur logging_steps : 100/5 = 20 points/seed — borderline pour intra-seed DM (≥22).\")\n", + "print(\"Pour garantir l'informatif, ajuster à logging_steps=4 (100/4 = 25 points/seed).\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-load-model-header", + "metadata": { + "papermill": { + "duration": 0.003504, + "end_time": "2026-08-25T01:49:37.132238", + "exception": false, + "start_time": "2026-08-25T01:49:37.128734", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 7. Chargement modèle Qwen3-1.7B ou Qwen3-2B + QLoRA 4-bit\n", + "\n", + "**Choix Qwen3-1.7B/2B vs Qwen3.5-0.8B** :\n", + "- **Plus de capacité** : 2.5× plus de paramètres pour absorber le signal RLVR\n", + "- **QLoRA 4-bit** : NF4 + double quant + bf16 compute — ~1.0 Go VRAM base pour 1.7B (vs ~0.5 Go pour 0.8B)\n", + "- **LoRA r=8, alpha=16** sur q/k/v/o/gate/up/down_proj — couverture complète MLP+attention\n", + "- **VRAM cible** : pic total < 8 Go (RTX 3070) ou < 14 Go (RTX 3080 Ti 16GB) — à mesurer en runtime\n", + "\n", + "**Pourquoi Qwen3 et pas Qwen3.5** :\n", + "- Qwen3 (1.7B, 2B) et Qwen3.5 (0.8B) coexistent dans HF\n", + "- Qwen3-1.7B est plus largement téléchargé (modèle de base) que Qwen3.5 équivalent\n", + "- Pour PT-11c, on cible le **gap** mesuré dans #12653 (groundage) : aucun GRPO exécuté au-delà de 0.8B dans le dépôt" + ] + }, + { + "cell_type": "code", + "execution_count": 10, + "id": "pt11c-load-model", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:37.141411Z", + "iopub.status.busy": "2026-08-25T01:49:37.141411Z", + "iopub.status.idle": "2026-08-25T01:49:37.145916Z", + "shell.execute_reply": "2026-08-25T01:49:37.145412Z" + }, + "papermill": { + "duration": 0.010674, + "end_time": "2026-08-25T01:49:37.146921", + "exception": false, + "start_time": "2026-08-25T01:49:37.136247", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Modèle cible : Qwen3-1.7B\n", + " Architecture : transformer LLM dense, vocab 151936 (Qwen3)\n", + " License : Apache 2.0\n", + " VRAM estimée (4-bit NF4) : ~1.0 Go (1.7B) ou ~1.2 Go (2B)\n", + " Loader : AutoModelForCausalLM + BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type='nf4')\n" + ] + } + ], + "source": [ + "import os\n", + "MODEL_NAME_1_7B = \"Qwen/Qwen3-1.7B\" # Cran au-dessus, modèle de base\n", + "MODEL_NAME_2B = \"Qwen/Qwen3-2B\" # Cran au-dessus (plus gros)\n", + "MODEL_NAME = MODEL_NAME_1_7B if MODEL_SIZE == \"1.7B\" else MODEL_NAME_2B\n", + "\n", + "_model_basename = os.path.basename(MODEL_NAME)\n", + "print(f\"Modèle cible : {_model_basename}\")\n", + "print(f\" Architecture : transformer LLM dense, vocab 151936 (Qwen3)\")\n", + "print(f\" License : Apache 2.0\")\n", + "print(f\" VRAM estimée (4-bit NF4) : ~1.0 Go (1.7B) ou ~1.2 Go (2B)\")\n", + "print(f\" Loader : AutoModelForCausalLM + BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type='nf4')\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-training-header", + "metadata": { + "papermill": { + "duration": 0.004509, + "end_time": "2026-08-25T01:49:37.155431", + "exception": false, + "start_time": "2026-08-25T01:49:37.150922", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 8. Training RLVR multi-seed — boucle 4 seeds × 100 steps\n", + "\n", + "**C'est le cœur de PT-11c.** On lance successivement le trainer GRPO pour chaque seed ∈ {0, 1, 7, 42}, en capturant `trainer.state.log_history` (reward par step) et `REWARD_ALERTS` (détecteur reward hacking par seed). Les données sont sauvegardées en JSONL dans `pt11c_per_seed_metrics.jsonl` pour l'analyse DM ci-dessous.\n", + "\n", + "**Wallclock estimé** : ~60 min/seed × 4 seeds = ~4h sur RTX 3080 Ti 16GB. Cycles c.558.\n", + "\n", + "**Critère d'arrêt anticipé** :\n", + "- `loss` explose à > 10× valeur initiale → divergent, rollback\n", + "- `reward` stagne à > 50 steps sans progression → convergence prématurée, signe de mode collapse\n", + "- OOM (out of memory) sur pic VRAM → réduire `per_device_train_batch_size` à 4 (toujours divisible par 8 grâce à `gradient_accumulation_steps=2`)" + ] + }, + { + "cell_type": "code", + "execution_count": 11, + "id": "pt11c-training", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:37.164320Z", + "iopub.status.busy": "2026-08-25T01:49:37.163431Z", + "iopub.status.idle": "2026-08-25T01:49:37.176154Z", + "shell.execute_reply": "2026-08-25T01:49:37.175649Z" + }, + "papermill": { + "duration": 0.018727, + "end_time": "2026-08-25T01:49:37.177159", + "exception": false, + "start_time": "2026-08-25T01:49:37.158432", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Mode CPU-safe : LOAD_MODEL_AND_TRAIN=False, pas de CUDA, ou pas de JSONL.\n", + "Pour exécuter : passer LOAD_MODEL_AND_TRAIN=True et GPU avec >= 4 Go VRAM.\n" + ] + } + ], + "source": [ + "if LOAD_MODEL_AND_TRAIN and CUDA_AVAILABLE:\n", + " import json as _json\n", + " import time as _time\n", + " from pathlib import Path as _Path\n", + " from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig\n", + " from trl import GRPOTrainer, GRPOConfig\n", + " from peft import LoraConfig, TaskType\n", + " import torch as _torch\n", + "\n", + " bnb_config = BitsAndBytesConfig(\n", + " load_in_4bit=True,\n", + " bnb_4bit_quant_type=\"nf4\",\n", + " bnb_4bit_compute_dtype=_torch.bfloat16,\n", + " bnb_4bit_use_double_quant=True,\n", + " )\n", + " lora_config = LoraConfig(\n", + " r=8,\n", + " lora_alpha=16,\n", + " lora_dropout=0.05,\n", + " bias=\"none\",\n", + " task_type=TaskType.CAUSAL_LM,\n", + " target_modules=[\"q_proj\", \"k_proj\", \"v_proj\", \"o_proj\", \"gate_proj\", \"up_proj\", \"down_proj\"],\n", + " )\n", + "\n", + " per_seed_metrics = []\n", + " metrics_path = _Path(OUTPUT_DIR) / \"pt11c_per_seed_metrics.jsonl\"\n", + "\n", + " for seed_idx, seed in enumerate(SEEDS):\n", + " print(\"=\" * 70)\n", + " print(f\" SEED {seed} ({seed_idx + 1}/{len(SEEDS)}) \")\n", + " print(\"=\" * 70)\n", + " REWARD_ALERTS.clear()\n", + " _GROUP_STATS.clear()\n", + "\n", + " t_load_start = _time.perf_counter()\n", + " base_model = AutoModelForCausalLM.from_pretrained(\n", + " MODEL_NAME,\n", + " quantization_config=bnb_config,\n", + " device_map=\"auto\",\n", + " )\n", + " tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)\n", + " if tokenizer.pad_token is None:\n", + " tokenizer.pad_token = tokenizer.eos_token\n", + " t_load = _time.perf_counter() - t_load_start\n", + " print(f\"Modèle chargé en {t_load:.1f} s\")\n", + "\n", + " # Mesure VRAM pic après chargement\n", + " vram_allocated_gb = _torch.cuda.memory_allocated(0) / 1e9\n", + " vram_reserved_gb = _torch.cuda.memory_reserved(0) / 1e9\n", + " print(f\" VRAM après load : {vram_allocated_gb:.2f} Go alloués, {vram_reserved_gb:.2f} Go réservés\")\n", + "\n", + " cfg_dict = dict(GRPO_CONFIG_DICT)\n", + " cfg_dict[\"seed\"] = seed\n", + " cfg_dict[\"output_dir\"] = f\"{OUTPUT_DIR}/seed_{seed}\"\n", + " grpo_config = GRPOConfig(**cfg_dict)\n", + "\n", + " trainer = GRPOTrainer(\n", + " model=base_model,\n", + " args=grpo_config,\n", + " processing_class=tokenizer,\n", + " train_dataset=dataset_rlvr,\n", + " reward_funcs=[rlvr_reward_func],\n", + " peft_config=lora_config,\n", + " )\n", + " print(f\"GRPOTrainer initialisé pour seed {seed}\")\n", + "\n", + " t_train_start = _time.perf_counter()\n", + " train_result = trainer.train()\n", + " t_train = _time.perf_counter() - t_train_start\n", + " print(f\"Training seed={seed} terminé en {t_train:.1f} s ({t_train/60:.1f} min)\")\n", + " print(f\" Loss finale : {train_result.training_loss:.4f}\")\n", + " print(f\" Alertes rewardspy : {len(REWARD_ALERTS)}\")\n", + "\n", + " vram_peak_gb = _torch.cuda.max_memory_allocated(0) / 1e9\n", + " print(f\" VRAM pic training : {vram_peak_gb:.2f} Go\")\n", + "\n", + " log_history = trainer.state.log_history if hasattr(trainer, 'state') else []\n", + " reward_curve = []\n", + " for entry in log_history:\n", + " step = entry.get('step')\n", + " reward = entry.get('reward')\n", + " if step is not None and reward is not None:\n", + " reward_curve.append((int(step), float(reward)))\n", + " print(f\" Reward curve : {len(reward_curve)} points\")\n", + "\n", + " _tot_inf = sum(ni for ni, _ in _GROUP_STATS)\n", + " _tot_grp = sum(nt for _, nt in _GROUP_STATS)\n", + " informative_fraction = (_tot_inf / _tot_grp) if _tot_grp > 0 else 0.0\n", + " print(f\" Groupes informatifs : {_tot_inf}/{_tot_grp} = {informative_fraction:.1%} (#10603)\")\n", + "\n", + " del trainer, base_model\n", + " _torch.cuda.empty_cache()\n", + " _torch.cuda.reset_peak_memory_stats(0)\n", + "\n", + " per_seed_metrics.append({\n", + " \"seed\": int(seed),\n", + " \"model_size\": MODEL_SIZE,\n", + " \"t_train_s\": float(t_train),\n", + " \"t_load_s\": float(t_load),\n", + " \"vram_peak_gb\": float(vram_peak_gb),\n", + " \"vram_allocated_gb\": float(vram_allocated_gb),\n", + " \"training_loss\": float(train_result.training_loss),\n", + " \"informative_fraction\": float(informative_fraction),\n", + " \"n_groups_total\": int(_tot_grp),\n", + " \"n_groups_informative\": int(_tot_inf),\n", + " \"reward_curve\": reward_curve,\n", + " \"alerts\": list(REWARD_ALERTS),\n", + " })\n", + " with open(metrics_path, 'a') as f:\n", + " f.write(_json.dumps(per_seed_metrics[-1]) + '\\n')\n", + " print(f\" Metrics persistés -> {metrics_path}\")\n", + " print()\n", + "\n", + " print(\"=\" * 70)\n", + " print(f\" TOUS LES SEEDS TERMINÉS ({len(SEEDS)} seeds) \")\n", + " print(\"=\" * 70)\n", + " print(f\"Metrics consolidés dans {metrics_path}\")\n", + "else:\n", + " metrics_path = Path(OUTPUT_DIR) / \"pt11c_per_seed_metrics.jsonl\"\n", + " if metrics_path.exists():\n", + " import json as _json_load\n", + " per_seed_metrics = []\n", + " with open(metrics_path) as _f:\n", + " for _line in _f:\n", + " if _line.strip():\n", + " per_seed_metrics.append(_json_load.loads(_line))\n", + " print(f\"Mode CPU-safe : training skip, mais JSONL détecté.\")\n", + " print(f\"Chargé {len(per_seed_metrics)} seeds depuis {metrics_path}\")\n", + " for _m in per_seed_metrics:\n", + " print(f\" seed={_m['seed']}: t_train={_m.get('t_train_s', 0):.0f}s, \"\n", + " f\"vram_peak={_m.get('vram_peak_gb', 0):.2f}Go, \"\n", + " f\"loss={_m.get('training_loss', 0):.4f}, \"\n", + " f\"n_pts={len(_m.get('reward_curve', []))}, \"\n", + " f\"alerts={len(_m.get('alerts', []))}\")\n", + " print(\"Cellules d'analyse peuvent être ré-exécutées sur ces données.\")\n", + " else:\n", + " print(\"Mode CPU-safe : LOAD_MODEL_AND_TRAIN=False, pas de CUDA, ou pas de JSONL.\")\n", + " print(\"Pour exécuter : passer LOAD_MODEL_AND_TRAIN=True et GPU avec >= 4 Go VRAM.\")\n", + " per_seed_metrics = []" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-curves-header", + "metadata": { + "papermill": { + "duration": 0.004714, + "end_time": "2026-08-25T01:49:37.188330", + "exception": false, + "start_time": "2026-08-25T01:49:37.183616", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 9. Courbes de reward par seed (overlay)\n", + "\n", + "Trace les 4 courbes de reward sur le même graphe pour voir la dispersion inter-seed. Une convergence stable = les 4 courbes montent et convergent ; une variance forte = signaux divergents entre seeds (bonhart détecteur de stochasticité réelle).\n", + "\n", + "**Lecture attendue pour PT-11c vs PT-11b** :\n", + "- PT-11b (0.8B, 2 seeds × 20 steps) : courbes plates, MECANISME_REPRO\n", + "- PT-11c (1.7B-2B, 4 seeds × 100 steps) : si BEATS, courbes ascendantes cross-seed ; si MECANISME_REPRO, courbes plates mais reproductibles" + ] + }, + { + "cell_type": "code", + "execution_count": 12, + "id": "pt11c-curves", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:37.197923Z", + "iopub.status.busy": "2026-08-25T01:49:37.197401Z", + "iopub.status.idle": "2026-08-25T01:49:37.520255Z", + "shell.execute_reply": "2026-08-25T01:49:37.520255Z" + }, + "papermill": { + "duration": 0.329243, + "end_time": "2026-08-25T01:49:37.521260", + "exception": false, + "start_time": "2026-08-25T01:49:37.192017", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Skip plot : pas de données de training (LOAD_MODEL_AND_TRAIN=False)\n" + ] + } + ], + "source": [ + "import matplotlib.pyplot as plt\n", + "from pathlib import Path\n", + "\n", + "if LOAD_MODEL_AND_TRAIN and CUDA_AVAILABLE and 'per_seed_metrics' in dir() and per_seed_metrics:\n", + " plt.figure(figsize=(10, 6))\n", + " colors = ['#2B5C8C', '#C44E52', '#55A868', '#8172B3']\n", + "\n", + " for i, m in enumerate(per_seed_metrics):\n", + " curve = m['reward_curve']\n", + " if not curve:\n", + " continue\n", + " steps, vals = zip(*curve)\n", + " plt.plot(steps, vals, marker='o', linewidth=1.5, alpha=0.85,\n", + " color=colors[i % len(colors)], label=f\"seed {m['seed']}\", markersize=3)\n", + "\n", + " plt.xlabel('Step')\n", + " plt.ylabel('Reward (outcome verifier)')\n", + " plt.title(f'PT-11c RLVR — Qwen3-{MODEL_SIZE} x {len(per_seed_metrics)} seeds (100 steps each)')\n", + " plt.grid(True, alpha=0.3)\n", + " plt.legend(loc='lower right')\n", + "\n", + " png_path = Path(\"MyIA.AI.Notebooks/GenAI/PostTraining/pt11c_reward_curves.png\")\n", + " png_path.parent.mkdir(parents=True, exist_ok=True)\n", + " plt.savefig(png_path, dpi=100, bbox_inches='tight')\n", + " print(f\"Figure sauvegardée : {png_path}\")\n", + " plt.show()\n", + "else:\n", + " print(\"Skip plot : pas de données de training (LOAD_MODEL_AND_TRAIN=False)\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-analysis-header", + "metadata": { + "papermill": { + "duration": 0.00298, + "end_time": "2026-08-25T01:49:37.528258", + "exception": false, + "start_time": "2026-08-25T01:49:37.525278", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 10. Analyse cross-seed — edge + Diebold-Mariano (linear)\n", + "\n", + "**Cœur du verdict.** Trois tests :\n", + "\n", + "1. **`edge_sigma` cross-seed** : `mean(seeds_mean_rewards) / std_dev_inter_seed`. Edge ≥ 2σ = signal au-dessus du bruit de seeds.\n", + "2. **`Diebold-Mariano` (DM)** sur la série de rewards par step, `loss_fn='linear'` (préserve le signe — mse/mae sont symétriques et rendent `dm_stat` bit-identique pour `e` et `-e`). `dm_p_median < 0.05` = significativité.\n", + "3. **Cohérence des deux** : BEATS uniquement si les **deux** conditions sont satisfaites (règle C du harness)." + ] + }, + { + "cell_type": "code", + "execution_count": 13, + "id": "pt11c-analysis", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:37.536478Z", + "iopub.status.busy": "2026-08-25T01:49:37.536478Z", + "iopub.status.idle": "2026-08-25T01:49:37.549548Z", + "shell.execute_reply": "2026-08-25T01:49:37.549548Z" + }, + "papermill": { + "duration": 0.01729, + "end_time": "2026-08-25T01:49:37.549548", + "exception": false, + "start_time": "2026-08-25T01:49:37.532258", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Skip DM analysis : pas de données de training (LOAD_MODEL_AND_TRAIN=False)\n" + ] + } + ], + "source": [ + "import sys as _sys\n", + "from pathlib import Path as _Path\n", + "_sys.path.insert(0, str(_Path(\"MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/scripts\").resolve()))\n", + "\n", + "import numpy as _np\n", + "from dm_test import diebold_mariano_test\n", + "\n", + "if 'per_seed_metrics' in dir() and per_seed_metrics:\n", + " # 1. Edge cross-seed\n", + " seed_mean_rewards = []\n", + " for m in per_seed_metrics:\n", + " if m['reward_curve']:\n", + " mean_r = _np.mean([v for _, v in m['reward_curve']])\n", + " seed_mean_rewards.append((m['seed'], mean_r))\n", + "\n", + " seeds_arr = _np.array([s for s, _ in seed_mean_rewards])\n", + " means_arr = _np.array([r for _, r in seed_mean_rewards])\n", + " overall_mean = _np.mean(means_arr)\n", + " inter_seed_std = _np.std(means_arr, ddof=1) if len(means_arr) > 1 else 0.0\n", + " edge_sigma = overall_mean / inter_seed_std if inter_seed_std > 1e-10 else float('inf')\n", + "\n", + " print(f\"Edge cross-seed : mean={overall_mean:.4f}, std={inter_seed_std:.4f}, edge={edge_sigma:.2f} sigma\")\n", + " for s, r in seed_mean_rewards:\n", + " print(f\" seed {s}: mean reward = {r:.4f}\")\n", + "\n", + " # 2. DM-vs-null baseline (info, pas décideur)\n", + " min_len = min(len(m['reward_curve']) for m in per_seed_metrics if m['reward_curve'])\n", + " print(f\"\\nDM setup : {len(per_seed_metrics)} seeds, {min_len} steps/seed, total obs={min_len * len(per_seed_metrics)}\")\n", + "\n", + " errors_model = _np.concatenate([\n", + " _np.array([v for _, v in m['reward_curve'][:min_len]])\n", + " for m in per_seed_metrics if m['reward_curve']\n", + " ])\n", + " errors_baseline = _np.zeros_like(errors_model)\n", + "\n", + " dm_pooled = diebold_mariano_test(errors_model, errors_baseline, loss_fn='linear')\n", + " print(f\"DM pooled (n={len(errors_model)}, linear) :\")\n", + " print(f\" dm_stat = {dm_pooled.dm_statistic:.4f}\")\n", + " print(f\" p_value = {dm_pooled.p_value:.6f}\")\n", + " print(f\" mean_loss_diff = {dm_pooled.mean_loss_diff:.4f}\")\n", + "\n", + " # DM per-seed (median p_value)\n", + " dm_per_seed = []\n", + " for m in per_seed_metrics:\n", + " if not m['reward_curve']:\n", + " continue\n", + " em = _np.array([v for _, v in m['reward_curve'][:min_len]])\n", + " eb = _np.zeros_like(em)\n", + " r = diebold_mariano_test(em, eb, loss_fn='linear')\n", + " dm_per_seed.append((m['seed'], r))\n", + " if dm_per_seed:\n", + " dm_p_median = _np.median([r.p_value for _, r in dm_per_seed])\n", + " print(f\"\\nDM per-seed (median p, n_seeds={len(dm_per_seed)}) :\")\n", + " for s, r in dm_per_seed:\n", + " print(f\" seed {s}: dm_stat={r.dm_statistic:.4f}, p={r.p_value:.6f}\")\n", + " print(f\" dm_p_median = {dm_p_median:.6f}\")\n", + " else:\n", + " dm_p_median = 1.0\n", + "\n", + " # 3. INTRA-SEED DM (décideur) : pre 20% vs post 20% par seed\n", + " intra_dm_per_seed = []\n", + " for m in per_seed_metrics:\n", + " rc = m.get('reward_curve', [])\n", + " if not rc or len(rc) < 22:\n", + " continue\n", + " vals = [v for _, v in rc]\n", + " n = len(vals)\n", + " cut = max(10, n // 5)\n", + " pre = _np.array(vals[:cut])\n", + " post = _np.array(vals[-cut:])\n", + " e_model = post - pre.mean()\n", + " e_baseline = pre - pre.mean()\n", + " r = diebold_mariano_test(e_model, e_baseline, loss_fn='linear')\n", + " intra_dm_per_seed.append((m['seed'], float(pre.mean()), float(post.mean()),\n", + " float(r.dm_statistic), float(r.p_value)))\n", + "\n", + " if intra_dm_per_seed:\n", + " intra_p_median = _np.median([r[4] for r in intra_dm_per_seed])\n", + " intra_mean_pre = _np.mean([r[1] for r in intra_dm_per_seed])\n", + " intra_mean_post = _np.mean([r[2] for r in intra_dm_per_seed])\n", + " intra_delta = intra_mean_post - intra_mean_pre\n", + " print(f\"\\nINTRA-SEED (pre 20% vs post 20%, par seed, n_seeds={len(intra_dm_per_seed)}) :\")\n", + " for s, mp, mq, ds, p in intra_dm_per_seed:\n", + " print(f\" seed {s}: pre={mp:.3f} post={mq:.3f} delta={mq-mp:+.3f} dm_stat={ds:.4f} p={p:.6f}\")\n", + " print(f\" intra_mean_pre = {intra_mean_pre:.4f}\")\n", + " print(f\" intra_mean_post = {intra_mean_post:.4f}\")\n", + " print(f\" intra_delta = {intra_delta:+.4f}\")\n", + " print(f\" intra_p_median = {intra_p_median:.6f}\")\n", + " else:\n", + " intra_p_median = 1.0\n", + " intra_delta = 0.0\n", + " print(\"\\nINTRA-SEED : 0 seed avec >=22 pts — non exécutable (logging_steps=5 donne 20 pts/seed).\")\n", + " print(\"Note : avec max_steps=100 et logging_steps=5, on a 20 points/seed (borderline).\")\n", + " print(\"Pour rendre l'intra-seed DM pleinement informatif, ajuster logging_steps=4 (-> 25 pts/seed).\")\n", + "\n", + " # 4. Verdict honnête — 4 classes :\n", + " edge_ok = edge_sigma >= 2.0\n", + " intra_ok = (intra_p_median < 0.05) and (intra_delta > 0)\n", + " if edge_ok and intra_ok:\n", + " verdict = \"BEATS\"\n", + " elif edge_ok and not intra_ok:\n", + " verdict = \"MECANISME_REPRO\"\n", + " elif (not edge_ok) and (not intra_ok):\n", + " verdict = \"NO BEATS\"\n", + " else:\n", + " verdict = \"INCONCLUSIVE\"\n", + " print(f\"\\nVERDICT FINAL : {verdict}\")\n", + " print(f\" edge >= 2 sigma cross-seed : {edge_ok} (edge_sigma={edge_sigma:.2f})\")\n", + " print(f\" intra-seed delta>0 et DM p<0.05 : {intra_ok} \"\n", + " f\"(intra_delta={intra_delta:+.4f}, intra_p_median={intra_p_median:.6f})\")\n", + " print(f\" (info) DM-vs-null baseline : dm_p_median={dm_p_median:.6f}\")\n", + "\n", + " PT11C_DM_RESULTS = {\n", + " \"edge_sigma\": float(edge_sigma),\n", + " \"dm_p_median\": float(dm_p_median),\n", + " \"dm_pooled_p\": float(dm_pooled.p_value),\n", + " \"intra_p_median\": float(intra_p_median),\n", + " \"intra_delta\": float(intra_delta),\n", + " \"intra_n_seeds\": len(intra_dm_per_seed),\n", + " \"verdict\": verdict,\n", + " \"n_seeds\": len(per_seed_metrics),\n", + " \"n_steps_per_seed\": min_len,\n", + " \"model_size\": MODEL_SIZE,\n", + " }\n", + "else:\n", + " print(\"Skip DM analysis : pas de données de training (LOAD_MODEL_AND_TRAIN=False)\")\n", + " PT11C_DM_RESULTS = {\"verdict\": \"INCONCLUSIVE_CPU_SAFE\", \"edge_sigma\": 0.0, \"dm_p_median\": 1.0,\n", + " \"intra_p_median\": 1.0, \"intra_delta\": 0.0, \"intra_n_seeds\": 0,\n", + " \"n_seeds\": 0, \"n_steps_per_seed\": 0, \"model_size\": MODEL_SIZE}" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-verdict-header", + "metadata": { + "papermill": { + "duration": 0.003792, + "end_time": "2026-08-25T01:49:37.557396", + "exception": false, + "start_time": "2026-08-25T01:49:37.553604", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 11. Verdict PT-11c — opposition directe à PT-11b\n", + "\n", + "**Acceptance finale** (falsifiable, règle §C du harness) :\n", + "\n", + "- **`BEATS`** : edge ≥ 2σ cross-seed ET (intra-seed DM p<0.05 `loss_fn='linear'` ET delta > 0)\n", + "- **`MECANISME_REPRO`** : edge ≥ 2σ mais intra-seed non significatif (reproductible, pas d'amélioration)\n", + "- **`NO BEATS`** : aucun des deux\n", + "- **`INCONCLUSIVE`** : l'un seul\n", + "\n", + "**Comparaison vs PT-11b** :\n", + "| Aspect | PT-11b (0.8B) | PT-11c (1.7B/2B) |\n", + "|---|---|---|\n", + "| Seeds × steps | 2 × 20 | 4 × 100 |\n", + "| Total obs | 40 | 400 |\n", + "| Edge cross-seed | 8.19σ | (à mesurer) |\n", + "| Intra-seed DM | non informatif (20<22) | **informatif (≥22)** |\n", + "| Verdict attendu | MECANISME_REPRO | BEATS / MECANISME_REPRO / NO BEATS |\n", + "\n", + "**Honnêteté fondamentale** : si PT-11c livre un autre `MECANISME_REPRO`, ce n'est **PAS** un échec — c'est une **convergence empirique** : GRPO ne crée pas la qualité à 0.8B ni à 1.7B-2B (cf convergence empirique PT-11b avec JohnEnev Part 3/4). Mais avant de conclure, il faut le **mesurer** : c'est le rôle de PT-11c." + ] + }, + { + "cell_type": "code", + "execution_count": 14, + "id": "pt11c-verdict", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:37.564965Z", + "iopub.status.busy": "2026-08-25T01:49:37.564965Z", + "iopub.status.idle": "2026-08-25T01:49:37.568841Z", + "shell.execute_reply": "2026-08-25T01:49:37.568841Z" + }, + "papermill": { + "duration": 0.00945, + "end_time": "2026-08-25T01:49:37.569845", + "exception": false, + "start_time": "2026-08-25T01:49:37.560395", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "======================================================================\n", + " VERDICT PT-11c — RLVR multi-seed sur Qwen3-1.7B/2B \n", + "======================================================================\n", + "Modèle : Qwen3-1.7B\n", + "Seeds : 0\n", + "Steps/seed : 0\n", + "edge_sigma : 0.00 (cross-seed)\n", + "dm_p_median : 1.000000 (DM-vs-null, info)\n", + "dm_pooled_p : 1.000000 (DM-vs-null pooled, info)\n", + "intra_p_median : 1.000000 (intra-seed DM, decideur)\n", + "intra_delta : +0.0000 (mean(post) - mean(pre))\n", + "intra_n_seeds : 0 (seeds avec >=22 pts)\n", + "Verdict : INCONCLUSIVE_CPU_SAFE\n", + "======================================================================\n", + "FIN VERDICT\n", + "======================================================================\n" + ] + } + ], + "source": [ + "print(\"=\" * 70)\n", + "print(\" VERDICT PT-11c — RLVR multi-seed sur Qwen3-1.7B/2B \")\n", + "print(\"=\" * 70)\n", + "\n", + "if 'PT11C_DM_RESULTS' in dir():\n", + " r = PT11C_DM_RESULTS\n", + " print(f\"Modèle : Qwen3-{r.get('model_size', 'N/A')}\")\n", + " print(f\"Seeds : {r.get('n_seeds', 'N/A')}\")\n", + " print(f\"Steps/seed : {r.get('n_steps_per_seed', 'N/A')}\")\n", + " print(f\"edge_sigma : {r.get('edge_sigma', 0.0):.2f} (cross-seed)\")\n", + " print(f\"dm_p_median : {r.get('dm_p_median', 1.0):.6f} (DM-vs-null, info)\")\n", + " print(f\"dm_pooled_p : {r.get('dm_pooled_p', 1.0):.6f} (DM-vs-null pooled, info)\")\n", + " print(f\"intra_p_median : {r.get('intra_p_median', 1.0):.6f} (intra-seed DM, decideur)\")\n", + " print(f\"intra_delta : {r.get('intra_delta', 0.0):+.4f} (mean(post) - mean(pre))\")\n", + " print(f\"intra_n_seeds : {r.get('intra_n_seeds', 0)} (seeds avec >=22 pts)\")\n", + " print(f\"Verdict : {r.get('verdict', 'INCONCLUSIVE')}\")\n", + "else:\n", + " print(\"INCONCLUSIVE_DATA_UNAVAILABLE — re-execution pendante\")\n", + " print(\"Aucune cellule d'analyse n'a été exécutée dans cette session.\")\n", + "\n", + "print(\"=\" * 70)\n", + "print(\"FIN VERDICT\")\n", + "print(\"=\" * 70)" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-exo3-header", + "metadata": { + "papermill": { + "duration": 0.004008, + "end_time": "2026-08-25T01:49:37.576852", + "exception": false, + "start_time": "2026-08-25T01:49:37.572844", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "### Exercice 3 : comparer PT-11c (1.7B) vs PT-11b (0.8B) sur le même JSONL\n", + "\n", + "Si PT-11c et PT-11b ont tous deux des JSONL disponibles, écrire une cellule qui :\n", + "1. Charge les deux JSONL\n", + "2. Calcule le mean reward global par modèle\n", + "3. Fait un DM test croisé (rewards PT-11c vs rewards PT-11b)\n", + "4. Conclut sur le **comparaison 0.8B vs 1.7B/2B** (BEATS 1.7B / BEATS 0.8B / EQUAL)\n", + "\n", + "**Objectif** : compléter l'acceptance #12653 (« comparaison honnête 0.8B vs cran supérieur avec verdict documenté »)." + ] + }, + { + "cell_type": "code", + "execution_count": 15, + "id": "pt11c-exo3", + "metadata": { + "execution": { + "iopub.execute_input": "2026-08-25T01:49:37.584008Z", + "iopub.status.busy": "2026-08-25T01:49:37.584008Z", + "iopub.status.idle": "2026-08-25T01:49:37.589227Z", + "shell.execute_reply": "2026-08-25T01:49:37.589227Z" + }, + "papermill": { + "duration": 0.010359, + "end_time": "2026-08-25T01:49:37.590232", + "exception": false, + "start_time": "2026-08-25T01:49:37.579873", + "status": "completed" + }, + "tags": [] + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Exercice à compléter : comparaison cross-cran PT-11b (0.8B) vs PT-11c (1.7B/2B)\n" + ] + } + ], + "source": [ + "def compare_11b_vs_11c(jsonl_11b_path: str, jsonl_11c_path: str) -> dict:\n", + " \"\"\"TODO etudiant : comparer mean reward PT-11b (0.8B) vs PT-11c (1.7B/2B).\"\"\"\n", + " # Étape 1 : charger les deux JSONL\n", + " # Étape 2 : calculer mean reward global par modèle\n", + " # Étape 3 : DM test croisé (loss_fn='linear')\n", + " # Étape 4 : verdict BEATS_1_7B / BEATS_0_8B / EQUAL / INCONCLUSIVE\n", + " return {\"verdict\": \"INCONCLUSIVE\", \"reason\": \"À implémenter\"}\n", + "\n", + "print(\"Exercice à compléter : comparaison cross-cran PT-11b (0.8B) vs PT-11c (1.7B/2B)\")" + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-bilan-header", + "metadata": { + "papermill": { + "duration": 0.003, + "end_time": "2026-08-25T01:49:37.597251", + "exception": false, + "start_time": "2026-08-25T01:49:37.594251", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## Bilan — RLVR sur le cran au-dessus : verdict mémoire + verdict BEATS\n", + "\n", + "PT-11b (0.8B, 2×20) a livré le verdict `MECANISME_REPRO` (edge cross-seed = 8.19σ, intra-seed DM non informatif). PT-11c (1.7B-2B, 4×100) pousse le protocole jusqu'au bout :\n", + "\n", + "1. **Verdict mémoire** : pic VRAM mesuré sur RTX 3080 Ti 16GB. Cible : tenir en < 14 Go (l'étage GPU moyen = RTX 3080 Ti). Si OOM : fallback RTX 3090 24GB ou A100 40GB.\n", + "2. **Verdict BEATS** : conjonction §C (edge ≥ 2σ ET intra-seed DM significatif). 4 seeds × 100 steps rendent l'intra-seed DM pleinement informatif (≥22 pts/seed vs 20 borderline).\n", + "3. **Comparaison 0.8B vs 1.7B/2B** : si BEATS, alors « le cran au-dessus est meilleur à budget steps égal ». Sinon, convergence empirique avec PT-11b : GRPO n'améliore pas la qualité sur ces tailles, quel que soit le cran.\n", + "\n", + "**Limites reconnues** :\n", + "- Si PT-11c livre `MECANISME_REPRO` (= GRPO reproductible mais pas d'amélioration), c'est une **convergence empirique** forte avec PT-11b et les mesures externes (JohnEnev Part 3/4). Le takeaway opérationnel : **investir dans le SFT d'abord, RL ensuite**.\n", + "- Si PT-11c livre `BEATS` (= GRPO améliore à 1.7B-2B), c'est un **signal nouveau** : le plafond RLVR pourrait être plus haut que les mesures à 0.8B ne le suggéraient. À creuser avec PT-12 (multi-step delayed credit).\n", + "- Si PT-11c livre `NO BEATS` ou `INCONCLUSIVE`, c'est un signal **négatif** : GRPO peut dégrader à 1.7B-2B, ce qui invaliderait partiellement l'hypothèse d'un plafond.\n", + "\n", + "**Honnêteté méthodologique** : tous les verdict sont également valides — un `MECANISME_REPRO` est une **donnée**, pas un échec. C'est la **stabilité du protocole** qui fait sa valeur : 4 seeds × 100 steps + conjonction §C = verdict falsifiable." + ] + }, + { + "cell_type": "markdown", + "id": "pt11c-transition-header", + "metadata": { + "papermill": { + "duration": 0.003003, + "end_time": "2026-08-25T01:49:37.604233", + "exception": false, + "start_time": "2026-08-25T01:49:37.601230", + "status": "completed" + }, + "tags": [] + }, + "source": [ + "## 12. Transition vers PT-12 — Multi-step delayed credit (cycle futur)\n", + "\n", + "PT-12 (déjà livré sur `myia-ai-01:CoursIA` GPU 2× RTX 4090 24GB) traite un **verdict adjacent** : GAE-λ sur env multi-step à crédit différé causal. PT-11c + PT-12 = couverture complète des estimateurs RL (REINFORCE/GRPO/RLOO/GAE-λ) sur petits modèles.\n", + "\n", + "PT-11c ferme la tranche **cran au-dessus 0.8B**. PT-13+ (futur) explorera le cran **7B+** sur RTX 3090 24GB avec QLoRA 4-bit et GRPO batch=4 (le cran 7B ne tient pas sur 16GB en G=8)." + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "Python 3", + "language": "python", + "name": "python3" + }, + "language_info": { + "codemirror_mode": { + "name": "ipython", + "version": 3 + }, + "file_extension": ".py", + "mimetype": "text/x-python", + "name": "python", + "nbconvert_exporter": "python", + "pygments_lexer": "ipython3", + "version": "3.11.9" + }, + "papermill": { + "default_parameters": {}, + "duration": 14.951227, + "end_time": "2026-08-25T01:49:38.926908", + "environment_variables": {}, + "exception": null, + "input_path": "PT_11c_grpo_qwen17_rlvr.ipynb", + "output_path": "PT_11c_grpo_qwen17_rlvr.ipynb", + "parameters": {}, + "start_time": "2026-08-25T01:49:23.975681", + "version": "2.6.0" + } + }, + "nbformat": 4, + "nbformat_minor": 5 +} \ No newline at end of file diff --git a/MyIA.AI.Notebooks/GenAI/PostTraining/README.md b/MyIA.AI.Notebooks/GenAI/PostTraining/README.md index 92a204819c..35a23b83a3 100644 --- a/MyIA.AI.Notebooks/GenAI/PostTraining/README.md +++ b/MyIA.AI.Notebooks/GenAI/PostTraining/README.md @@ -4,12 +4,12 @@ -> **Place dans GenAI** : cette série est le pendant *théorique et SOTA 2024-2025* de la série [FineTuning](../FineTuning/README.md). FineTuning couvre la boîte à outils pratique (LoRA, QLoRA, SFT, DPO, model merging) sur 5 notebooks exécutés ; PostTraining remonte la chaîne conceptuelle complète SFT → RLHF → DPO → GRPO → RLVR → **GAE** et reproduit les techniques récentes (Deepseek-R1) sur petits modèles, complétée par un notebook d'évaluation comparative, un détecteur de reward hacking, et un notebook d'implémentation from-scratch de la famille "no critic" (GRPO/RLOO/GAE) sur toy env CPU, par un **notebook multi-step à crédit différé causal** (PT-12 : les cinq estimateurs re-mesurés, GAE-λ devient discriminant, verdict BEATS 5/5 seeds — le "1-step collapse" était une propriété du banc), et de **deux notebooks appliqués Qwen3.5-0.8B + GRPO + reward vérifiable + rewardspy en ligne** (PT-11a Z3 CSP arithmétique + PT-11b SymPy arithmétique + Z3 N-queens, plus leur validation multi-seed) qui font sortir la série du toy env vers un vrai LLM, soit **14 notebooks** au total. Les deux se complèment : commencer par FineTuning pour la pratique, PostTraining pour la profondeur méthodologique. +> **Place dans GenAI** : cette série est le pendant *théorique et SOTA 2024-2025* de la série [FineTuning](../FineTuning/README.md). FineTuning couvre la boîte à outils pratique (LoRA, QLoRA, SFT, DPO, model merging) sur 5 notebooks exécutés ; PostTraining remonte la chaîne conceptuelle complète SFT → RLHF → DPO → GRPO → RLVR → **GAE** et reproduit les techniques récentes (Deepseek-R1) sur petits modèles, complétée par un notebook d'évaluation comparative, un détecteur de reward hacking, et un notebook d'implémentation from-scratch de la famille "no critic" (GRPO/RLOO/GAE) sur toy env CPU, par un **notebook multi-step à crédit différé causal** (PT-12 : les cinq estimateurs re-mesurés, GAE-λ devient discriminant, verdict BEATS 5/5 seeds — le "1-step collapse" était une propriété du banc), et de **trois notebooks appliqués Qwen + GRPO + reward vérifiable + rewardspy en ligne** (PT-11a Z3 CSP arithmétique sur Qwen3.5-0.8B + PT-11b SymPy arithmétique + Z3 N-queens, plus leur validation multi-seed, **plus PT-11c sur le cran au-dessus Qwen3-1.7B/2B** qui qualifie l'étage GPU moyen 16 Go et oppose 0.8B vs 1.7B/2B à budget steps égal) qui font sortir la série du toy env vers un vrai LLM, soit **15 notebooks** au total. Les deux se complèment : commencer par FineTuning pour la pratique, PostTraining pour la profondeur méthodologique. Série pédagogique dédiée aux techniques de **post-training** des LMs ouverts : SFT, DPO, GRPO, RLVR. L'objectif est de comprendre pourquoi 2024-2025 marque une rupture pédagogique dans la façon dont les modèles de langue passent du pre-training brut à un assistant utile, et comment cette chaîne s'est simplifiée depuis la cascade RLHF historique jusqu'aux méthodes "direct" récentes. @@ -40,6 +40,7 @@ L'angle pédagogique est d'expliquer la **math du loss** avant le code pour chaq | PT-11a | `PT_11_grpo_qwen35_rlvr.ipynb` | GRPO + RLVR sur **vrai LLM** (Qwen3.5-0.8B QLoRA 4-bit) — reward vérifiable Z3 (CSP arithmétique), rewardspy **en ligne**, sortie du toy env | `trl.GRPOTrainer` + Z3 + `rewardspy.watch_trl` | Qwen3.5-0.8B (QLoRA 4-bit, GPU 8 Go) | See #10289 (#10302) | | PT-11b | `PT_11_grpo_qwen_rlvr_on_verifiers.ipynb` | GRPO + RLVR + rewardspy **en ligne** — la pile complète sur petit modèle Qwen3.5-0.8B QLoRA, **Tier 1 SymPy (arithmétique) + Tier 2 Z3 (N-queens)**, détecteur Goodhart live, 100 steps réels — variante à verifiers complémentaires (sibling de PT-11a) | `trl.GRPOTrainer` + `rewardspy.watch_trl` (integration trl 1.9.2) | Qwen3.5-0.8B (QLoRA 4-bit) | See #10289 | | PT-11b multi-seed | `PT_11b_multiseed_qwen35_4x100.ipynb` | RLVR **multi-seed** 4 seeds × 100 steps — reproductibilité du run PT-11b (Qwen3.5-0.8B QLoRA 4-bit), opposition au run mono-seed, métrique informative de groupe (`num_generations`, #10603), verdict MECANISME_REPRO | `trl.GRPOTrainer` + verifiers SymPy/Z3 | Qwen3.5-0.8B (QLoRA 4-bit) | See #10289 | +| PT-11c | `PT_11c_grpo_qwen17_rlvr.ipynb` | RLVR sur **cran au-dessus** (Qwen3-1.7B ou Qwen3-2B QLoRA 4-bit) — extension directe de PT-11b au cran supérieur, 4 seeds × 100 steps, **verdict mémoire** (pic VRAM RTX 3080 Ti 16GB) + **comparaison honnête** 0.8B vs 1.7B/2B à budget steps égal. Architecture compatible `trl.GRPOTrainer`, peft 0.20.0 (vs 0.13.2 de PT-11b). Reward SymPy Tier-1 + Z3 N-queens Tier-2 + informative-group métrique #10603. **Verdict attendu** : BEATS si GRPO améliore à 1.7B/2B ; MECANISME_REPRO si convergence empirique avec PT-11b. **Note env** : `rewardspy` (GitHub-only, absent de PyPI) peut être non installé — wrapper no-op gracieux | `trl.GRPOTrainer` + verifiers SymPy/Z3 + `rewardspy.watch_trl` (optionnel) | Qwen3-1.7B ou Qwen3-2B (QLoRA 4-bit, GPU 16 Go cible) | See #12653 | | PT-12 | `PT_12_multistep_delayed_credit.ipynb` | GAE-λ sur env **multi-step à crédit différé causal** (`count_ones`, terminal dépendant de toute la séquence) — re-mesure des 5 estimateurs REINFORCE/GRPO/RLOO/GAE-λ=0/GAE-λ=0.95, **5 seeds**, vérificateur terminal Z3 | `torch` from-scratch (no `trl`) + Z3 (RLVR) | MLP jouet acteur-critique (CPU, ~5.5k params) | See #1454 | > **Migration vers Qwen3.5 (complète sur les notebooks GPU).** Qwen2.5 est *superseded* par [Qwen3.5](https://huggingface.co/Qwen/Qwen3.5-0.8B) (modèle vision-langage unifié, multimodal). **PT-03 est migré** vers Qwen3.5-0.8B (#5078) : l'évaluation DPO est désormais une vraie *forward pass* (accuracy mesurée 40 % sur 10 préférences *held-out*, vs 50 % aléatoire — le DPO n'a pas convergé sur 50 exemples, verdict honnête documenté dans le notebook) qui remplace l'ancienne accuracy *hardcodée* à 72 %. **PT-02 est migré** vers Qwen3.5-0.8B (#10289, #10813) : le SFT s'exécute sur le vrai petit modèle SOTA (architecture hybride 18×linear_attn + 6×self_attn, LoRA ciblant les deux types + MLP) — verdict honnête documenté : sur 50 exemples (3 steps), le loss oscille (1.92→2.17→1.76) et la génération se dégrade (SFT sous-alimenté sur modèle déjà instruct-tuned, pas un bug). **PT-04 est migré** vers Qwen3.5-0.8B (#10817, re-exec #11442) : le GRPO s'exécute sur le vrai petit modèle SOTA (vrai GRPO GPU, QLoRA 4-bit, récompense basée sur le score). **PT-06 est migré** vers Qwen3.5-0.8B (#10819) : l'évaluation comparative porte désormais sur le modèle migré. **PT-05 est migré** vers Qwen3.5-0.8B (#10487 : re-ciblage + run GPU réel commité — delta honnête documenté 12.5 % → 0.0 %, et #10504 : reward branché sur le ground truth). Architecture Qwen3.5 multimodale (`Qwen3_5ForConditionalGeneration`) chargée via `AutoModelForImageTextToText` (et non `AutoModelForCausalLM`).