Skip to content

fix(notebook,#18124): research_long_short_harvest -- real VIX quintiles + measured Hurst + trained ML core replace placeholder - #18164

Merged
myia-ai-01 merged 4 commits into
mainfrom
feature/18124-researchexec-longsshort
Sep 29, 2026
Merged

myia-ai-01 merged 4 commits into
mainfrom
feature/18124-researchexec-longsshort

Conversation

@jsboige

@jsboige jsboige commented Sep 28, 2026

Copy link
Copy Markdown
Owner

Grain: MED/notebook-python -- lane myia-po-2026:CoursIA -- prev: #18163

Summary

SOTA repair of Research-Executor/research_long_short_harvest.ipynb (last of
the placeholder set in the #18124 audit). Was: QuantBook-only (unrunnable
locally), analysis cells un-executed, fake "BACKTEST RESULTS" placeholder.

  • Local research path: SPY + GLD + the real ^VIX index via yfinance,
    tz-normalized, notebook executes end-to-end locally.
  • VIX quintile -> 21d forward SPY return measured on the real index:
    Q1 +0.82% / Q2 +0.40% / Q3 +0.90% / Q4 +1.04% / Q5 +2.87% (n~600 per
    bucket) -- the volatility-risk-premium structure the strategy builds on,
    measured instead of asserted.
  • Hurst exponent measured (aggregated variance): SPY 0.009, GLD 0.020 --
    daily-scale series sit far below the H > 0.85 short-screen threshold;
    the notebook now says how selective that screen really is.
  • The promised RandomForestClassifier trained for real (11 VIX/SPY
    features, temporal 70/30 split): AUC 0.386, accuracy 61.2% < 69.8%
    majority baseline -- honest NO EDGE verdict written in, not hidden.
  • Full long-short engine (top-4 market-cap longs, weekly Hurst shorts,
    3-stage stops, margin) explicitly routed to QC Cloud.

Validation

  • Papermill: 9/9 cells, 0 errors, all execution_count set (C.2)
  • C.1: no raise NotImplementedError / assert False / 1/0 (verified)
  • Catalogue byte-identical to main (single .ipynb changed)
  • SOTA verdict: SOTA-OK -- real index data, measured statistics, trained
    model, honest out-of-sample verdict

See #18124. With this PR the tranche claimed by this lane is complete:
6 of the 10 audited notebooks repaired (fallback-banner trio + placeholder
trio); remaining 4 (non-executed-cells family: defensive 4 done in #18163,
macro_factor_rotation + any stragglers) to re-audit against the current state.

🤖 Generated with Claude Code

…es + measured Hurst + trained ML core replace placeholder

SOTA repair (last of the placeholder set in the #18124 audit). The notebook
was QuantBook-only with un-executed analysis cells and a fake backtest
placeholder at the end.

- Local research path: SPY + GLD + the real ^VIX index via yfinance
  (tz-normalized), so the whole notebook executes locally.
- VIX quintile -> next-21d SPY return table measured on the real index:
  Q1 +0.82% ... Q5 +2.87% (n~600 per bucket) -- the volatility-risk-premium
  structure the strategy builds on, measured instead of asserted.
- Hurst exponent measured (aggregated variance) on SPY 0.009 / GLD 0.020:
  daily-scale series sit far below the H > 0.85 short-screen threshold --
  informative about how selective that screen is.
- The promised RandomForestClassifier (11 VIX/SPY features) trained for real
  (temporal 70/30 split): AUC 0.386, accuracy 61.2% < 69.8% majority
  baseline -- honest NO EDGE verdict written in.
- The full long-short engine (top-4 market-cap longs, weekly Hurst shorts,
  3-stage stops, margin) explicitly routed to QC Cloud, not faked.

Executed 9/9 cells, 0 errors, all execution_count set (C.1/C.2).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

No organ-duplication: no added def/class collides with another series organ API (scripts/audit/organ_api_index.yaml).

Detector: python scripts/audit/detect_organ_duplication.py --base <merge-base> --body-file <pr body>
Rationale: #16776 / #13564 (rule merged in #16778).

@github-actions

Copy link
Copy Markdown
Contributor

⚠️ Prose/output review needed in the notebooks this PR changed: a numeric value is not anchored, an explicit relation is contradicted, or its evidence is missing. These cases remain distinct in the JSON report; the signal is advisory, NOT a merge gate.

Scope = notebooks CHANGED in this PR, not the whole corpus. Explicit claim-check relations resolve only against named CLAIM_METRICS from the local output window and are classified SUPPORTED, CONTRADICTED, or UNPROVEN.
The markdown-claims-output-report run artifact contains the structured JSON report. See python scripts/check_markdown_claims_output.py --help for re-running locally.
Detector rationale: c.290 / c.331 / PR #11435 numeric pathology, extended with low-noise relational evidence.

@github-actions

Copy link
Copy Markdown
Contributor

Notebook outputs-required (H.4 schema): PASS (every code cell carries an outputs: list)

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Golden-Set Execution (H.7 P3)

✅ 8/8 notebooks passed (certified reproducible)

Notebook Status Time
2.1-Workflow-ML.ipynb ✅ SUCCESS 7.8s
2.2-Descente-de-gradient.ipynb ✅ SUCCESS 9.1s
2.3-Regression-lineaire-logistique.ipynb ✅ SUCCESS 14.5s
2.4-Arbres-Forets-Ensembles.ipynb ✅ SUCCESS 10.8s
Search-01-StateSpace.ipynb ✅ SUCCESS 8.2s
SL-1-LogicalLearning.ipynb ✅ SUCCESS 9.0s
rl_4_multi_armed_bandits.ipynb ✅ SUCCESS 60.5s
GameTheory-04c-NashExistence-Python.ipynb ✅ SUCCESS 7.5s

Pinned lockfile: scripts/notebook_tools/golden_set.lock.txt (H.7 P3, axe A #4208)

@github-actions

Copy link
Copy Markdown
Contributor

Notebook PR Validation: PASS

  • Notebooks checked: 1
  • Code cells validated: 4
  • Result: All passed

Checks: H.1 (no errors), H.3 (execution_count), C.1 (no banned patterns)
Non-Python kernels (.NET/Lean): C.1 + errors only (execution_count advisory)
QuantConnect notebooks: C.1 + errors only (require QC Cloud for execution)

@github-actions github-actions Bot added the large-pr-no-review PR > seuil sans review (ni bot ni humaine) -- retire quand une review arrive (#11232) label Sep 28, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Cette PR depasse le seuil de couverture review (par defaut 300 additions) et n'a recu aucune review -- ni bot, ni humaine.

Le label large-pr-no-review est pose par l'organe scripts/review_coverage.py porte par l'issue #11232. Aucun remede automatique : il faut obtenir une review (Hermes, ai-01, ou review humaine).

Le label est retire au balayage suivant (quotidien) des qu'une review arrive -- dans reviews[] ou en commentaire de verdict -- ou que le diff passe sous le seuil. Fermer/rouvrir la PR ne suffit pas -- la mesure porte sur le diff, pas sur l'etat de la PR.

Seuil, historique et exceptions : cf. docs/reference/review-coverage-threshold.md.

@jsboige

jsboige commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

[ADJOINT PREFLIGHT]
schema: 1
lane: myia-po-2023:CoursIA
pr: 18164
head: 76bfc4b
complete: true
body: read
comments-reviewed: 6
reviews-reviewed: 0
threads-reviewed: 0
threads-unresolved: 0
surfaces-sha256: 38481c9871f18f9383e6875548e1f6aac9dc9edf160ff47bcce3b0477de1a508
diff-files: 1
diff-additions: 329
diff-deletions: 89
checks: latest-wins-green
b0: clear
scope: pass
domain: not-applicable
verdict: READY
[/ADJOINT PREFLIGHT]

Note de lecture (informatif) : advisory sticky markdown-claims-output-advisory present (valeur numerique a tracer dans l artifact du run) -- non bloquant par design, meme classe que #18162.

@myia-ai-01
myia-ai-01 merged commit b038051 into main Sep 29, 2026
88 of 89 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

large-pr-no-review PR > seuil sans review (ni bot ni humaine) -- retire quand une review arrive (#11232) variation-tag-prev-absent Tag Grain sans 'prev: <TIER>/<GENRE> #<PR>' (adjacence G-VAR-3 inevaluable)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants