Skip to content

feat(gametheory,#17528 arc A): Assistance Games 2026 — POLA (Provably Optimal Learning) exécuté dans GameTheory-15 #17776

Description

@jsboige

Contexte

EPIC umbrella #17528 (Russell & Norvig — habiter l'armature AIMA) arc A demande explicitement :

Cible : un résultat de 2024-2026 exécuté, et non plus seulement cité. Candidats : Provably Optimal Learning Algorithms for Assistance Games, Learning the Preferences of a Learning Agent, The Partially Observable Off-Switch Game, AssistanceZero.

Le grain précédent #17529 (livré PR #17648 MERGEABLE) a traité le seuil override_threshold=0.9 (MED confrontation aux sources). Ce grain-ci porte la cible DEEP 2026.

Cible

Provably Optimal Learning Algorithms for Assistance Games (Ananthakrishnan, Bedaywi, Jordan, Russell, Haghtalab — 2026, arXiv 2607.08012). Source archivée :

  • Chemin :
  • sha8 :
  • Identité vérifiée sur première page (Tell c.974 strict ★★★)

Thèse : dans un assistance game (robot + humain, utilités alignées mais préférences humain inconnues), il existe un algorithme d'apprentissage prouvablement optimal (regret O(√T)) qui apprend en ligne les préférences de l'humain tout en convergeant vers l'équilibre Stackelberg.

Acceptance

  1. Section §4.4 « Assistance Games 2026 » dans après §4.3 SUR (cellules 35-36).
  2. Implémentation Python (sans copier-coller de la lib) : Assistance Game = environnement POMDP-like avec reward paramétrée par θ (préférence humaine). Deux agents : Robot (apprend θ), Human (signale via action).
  3. Algorithme 2026 : POLA (Provably Optimal Learner for Assistance) — algorithme de bandit bayésien avec posterior sur θ et choix d'action = argmax Stackelberg Value of Information (SVOI).
  4. Validation multi-seed ≥4 (seeds 0/1/7/42/99), walk-forward 5-fold :
    • Métrique 1 : regret cumulé vs optimum en info parfaite — doit être O(√T)
    • Métrique 2 : convergence Stackelberg : probabilité de choix optimal au round T — doit tendre vers 1
  5. Comparaison baseline : Robot greedy (choisit action maximisant reward moyenne) + Robot random.
  6. Verdict honnête « BEATS / NO BEATS / INCONCLUSIVE » avec DM p-value < 0.05 sur la métrique regret (mse).
  7. Sortie : courbes regret cumulé + Stackelberg probabilité + tableau des 3 algorithmes × 5 seeds.
  8. Cellule interprétation : « ### Lecture du résultat POLA vs Greedy » — expliquer ce que la mesure prouve sur la discrimination des algorithmes.
  9. Citations biblio : DOI/arXiv vérifiés, arXiv IDs CITÉS ≠ fabrications (Tell c.c.c.d.G.1 ★★★★).
  10. Pas de regression §4.1-4.3 : la réparation c.846-c.847 (PR fix(gametheory,#17529): seuil Off-Switch Game aligne sur override_threshold=0.9 #17648) doit rester verte (cellule 32 « refuse en deca de ce seuil »).

Périmètre

Hors scope

Dépendances

  • (déjà installé) — pas nécessaire pour Assistance Games (jeu custom), mais vérifier présence.
  • du dossier GameTheory (bibliothèque interne des sections §1-§4).
  • numpy, matplotlib.

Tell fondateur à respecter

  • Tell c.c.c.d.G.1 ★★★★ : vérif first-hand arXiv IDs AVANT citation.
  • Tell c.c.c.d.974 ★★★ : relecture INVERSION CHRONOLOGIQUE si reprise de §4.1-§4.3.
  • Tell c.c.c.d.c813 ★★★ : nbformat cell.source newline — chaque ligne doit finir par sauf la dernière.
  • sota-not-workaround.md §F : Assistance Game 2026 = moteur SOTA réel, pas une reimplementation jouet.

Grain REPAIR/DEEP — hérite du genre notebook-python (CONTENU), tier DEEP.

Grain: DEEP/notebook-python — lane myia-po-2023:CoursIA-2 — prev: REPAIR/notebook-python #17648

See #17528 (arc A)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions