You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat(gametheory,#17528 arc A): Assistance Games 2026 — POLA (Provably Optimal Learning) exécuté dans GameTheory-15 #17776
Cible : un résultat de 2024-2026 exécuté, et non plus seulement cité. Candidats : Provably Optimal Learning Algorithms for Assistance Games, Learning the Preferences of a Learning Agent, The Partially Observable Off-Switch Game, AssistanceZero.
Le grain précédent #17529 (livré PR #17648 MERGEABLE) a traité le seuil override_threshold=0.9 (MED confrontation aux sources). Ce grain-ci porte la cible DEEP 2026.
Cible
Provably Optimal Learning Algorithms for Assistance Games (Ananthakrishnan, Bedaywi, Jordan, Russell, Haghtalab — 2026, arXiv 2607.08012). Source archivée :
Chemin :
sha8 :
Identité vérifiée sur première page (Tell c.974 strict ★★★)
Thèse : dans un assistance game (robot + humain, utilités alignées mais préférences humain inconnues), il existe un algorithme d'apprentissage prouvablement optimal (regret O(√T)) qui apprend en ligne les préférences de l'humain tout en convergeant vers l'équilibre Stackelberg.
Acceptance
Section §4.4 « Assistance Games 2026 » dans après §4.3 SUR (cellules 35-36).
Implémentation Python (sans copier-coller de la lib) : Assistance Game = environnement POMDP-like avec reward paramétrée par θ (préférence humaine). Deux agents : Robot (apprend θ), Human (signale via action).
Algorithme 2026 : POLA (Provably Optimal Learner for Assistance) — algorithme de bandit bayésien avec posterior sur θ et choix d'action = argmax Stackelberg Value of Information (SVOI).
Contexte
EPIC umbrella #17528 (Russell & Norvig — habiter l'armature AIMA) arc A demande explicitement :
Le grain précédent #17529 (livré PR #17648 MERGEABLE) a traité le seuil override_threshold=0.9 (MED confrontation aux sources). Ce grain-ci porte la cible DEEP 2026.
Cible
Provably Optimal Learning Algorithms for Assistance Games (Ananthakrishnan, Bedaywi, Jordan, Russell, Haghtalab — 2026, arXiv 2607.08012). Source archivée :
Thèse : dans un assistance game (robot + humain, utilités alignées mais préférences humain inconnues), il existe un algorithme d'apprentissage prouvablement optimal (regret O(√T)) qui apprend en ligne les préférences de l'humain tout en convergeant vers l'équilibre Stackelberg.
Acceptance
Périmètre
Hors scope
Dépendances
Tell fondateur à respecter
Grain REPAIR/DEEP — hérite du genre notebook-python (CONTENU), tier DEEP.
Grain: DEEP/notebook-python — lane myia-po-2023:CoursIA-2 — prev: REPAIR/notebook-python #17648
See #17528 (arc A)