diff --git a/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/REGISTRY.md b/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/REGISTRY.md index 3fc81bef79..7a796d63a4 100644 --- a/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/REGISTRY.md +++ b/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/REGISTRY.md @@ -15,6 +15,10 @@ Updated: 2026-08-24 — M15 LSTM-vol patch persistance biais + slice 2/2 dé-bia Updated: 2026-09-01 — PatchTST BTC log-RV revalidé contre HAR débiaisé train-only (#14081) : h=1 INCONCLUSIVE, h=5/h=10 NO BEATS ; var_ratio > 1 aux trois horizons Updated: 2026-09-02 — M16 HAR asymétrique BTC revalidé contre HAR débiaisé train-only (#1454) : h=1 INCONCLUSIVE, h=5/h=10 BEATS ; verdict brut 3/3 réfuté Updated: 2026-09-02 — M5 HMM regime-switching HAR, première entrée + revalidation hors biais (Epic #1454) : ETH h=1 **BEATS confirmé** (+8,7 % hors biais, 4/4 seeds, 4,1σ) ; BTC h=1 s'effondre de +7,0 % à +1,3 % (~82 % de l'edge était le biais de HAR) ; 4/6 NO BEATS +Updated: 2026-09-04 — M17 HAR-LJ-Asym BTC revalidé contre HAR débiaisé train-only (Epic #1454, lane myia-po-2026, c.951) : h=1 BEATS 4/4 confirmé (var_ratio 0,778 — gain de précision, pas un offset) ; h=5/h=10 INCONCLUSIVE 0/4 (var_ratio > 1) ; vs M12 4/4 BEATS aux trois horizons ; verdict BTC inchangé, base non-artéfactée par le biais HAR +Updated: 2026-09-04 — M17 HAR-LJ-Asym BTC REPAIR P0 c.953 (PR #14592, preflight po-2025 adjoint `msg-20260904T105224-z6f9d7` head 8167044f) : **calibration symétrique** sur LJ/HAR/M12 + `var(ddof=0)` + sanity `mse = bias²+var` à 1e-9 + `panel_hash` SHA256 sur fenêtre canonique 360 bars (BTC 4 seeds → 1 hash `86f36cb46f539c6d`) + naming `mse_har_raw`/`mse_har_debiased` (NaN si non-débiaisé). **Verdict révisé** : h=1 BEATS 4/4 (inchangé, var_ratio=0,778) ; h=5 INCONCLUSIVE 0/4 contre HAR **et** M12 (était 4/4 BEATS vs M12) ; h=10 BEATEN BY 0/4 contre HAR (MSE 0,464 > 0,366 HAR-débiaisé ; était INCONCLUSIVE 0/4). Le claim c.951 « 4/4 BEATS vs M12 aux trois horizons » ne tient plus sous calibration symétrique : la tête précédente de M17 sur M12 aux h=5/h=10 était portée par le gap de calibration, pas par un gain de précision. Detail dans `docs/M17_HAR_LJ_ASYM.md` section c.953. `panel_hashes_consistent=True`, OLS bit-identique par seed (DM-MSE p-values identiques 6 décimales). +Updated: 2026-09-05 — M17 HAR-LJ-Asym round-3 calibration (PR #14592, preflight po-2025 adjoint re-review head `b974f2721`, DM `msg-20260904T141944`) : **nombres c.953 ci-dessus SUPERSEDES** (signe du biais inversé `yhat - bias` → corrigé `yhat + bias`, application per-fold, M12 calibré `calibrate_bias=debias`, deux jambes HAR `mse_har_raw` != `mse_har_debiased`, `panel_hash` sur index+valeurs avec manifeste per (coin,horizon)). Détail dans `docs/M17_HAR_LJ_ASYM.md` section « Round-3 calibration (this PR) ». [M17 HAR-LJ-Asym BTC run] — pending live run post-merge; round-3 calibration implemented, code PR #14592. +Updated: 2026-09-05 — M17 HAR-LJ-Asym round-4 (PR #14592, adjoint re-review DM `msg-20260905T001520`, 3/6 PASS / 3/6 PARTIAL) : test OOS multi-fold discriminant `test_walk_forward_lj_asym_oos_target_invariance_multi_fold` (n_splits=3, biais per-fold **distincts**, invariance per-fold bit-identique rtol 1e-12, folds antérieurs inchangés, folds postérieurs = expanding-window retrain légitime asserté comme sensibilité, train-tail > 1.0 par fold) + provenance `bounds_train_test` (`{train_end_idx = n_splits·(n//(n_splits+1)), oos_start_idx = train_end+horizon, oos_end_idx}`) relayée par `_eval_one_coin` / `aggregate_verdicts` / manifeste (`bounds_per_coin_horizon`, `fc_lj_hash_per_fold` alignés sur `per_fold_bias`) ; placeholder `if False else None` supprimé. Tests 22 → 24 verts, suite 1194 passed / 0 failed. **[M17 HAR-LJ-Asym BTC run] LIVRÉ (round-4 code) — `python har_lj_asym.py --coins BTC-USD --skip-remote --debias --horizons 1 5 10 --seeds 0 7 42 99` en 467.9 s** : h=1 **BEATS 4/4** vs HAR et M12 (p<1e-6, mean_loss_diff<0, `_coherent_beats()` strict ✓) ; h=5/h=10 INCONCLUSIVE 0/4 (p>0.05, mean_loss_diff>0 ⇒ cohérent INCONCLUSIVE). Bornes effectives BTC : `train_end=1890` (5 folds × 378 jours), `n_oos=378-382`, `n_total=2272` jours. Bit-identity cross-seed OK (`per_fold_bias` et `fc_lj_hash_per_fold` identiques sur les 4 seeds, `bounds_consistent_across_seeds=True`). Précédent c.953 `h=1 BEATS p=0.839708` réfuté — sous round-3+4 calibration le verdict reste BEATS mais devient réellement significatif. Détail dans `docs/M17_HAR_LJ_ASYM.md` section « Live BTC run (concern b — this PR) ». Manifest `scripts/results/m17_har_lj_asym.json` régénéré ; meta `manifest_m17_har_lj_asym.json` mis à jour avec `concern_addressing` round-3 + round-4. Total checkpoints: 70 (20 legacy ARCHIVED + 50 panier baselines) diff --git a/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/docs/M17_HAR_LJ_ASYM.md b/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/docs/M17_HAR_LJ_ASYM.md index f575db0129..19db00b48c 100644 --- a/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/docs/M17_HAR_LJ_ASYM.md +++ b/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/docs/M17_HAR_LJ_ASYM.md @@ -135,3 +135,239 @@ The corrected sweep supersedes the bugged "60/60 BEATS" verdict (the bug made `horizon` a no-op via a contemporaneous target). Runtime: 261.5s for 84 combos (local, `--skip-remote` false). + +## Bias-debiased HAR baseline revalidation (BTC, c.951 → c.953) + +**Date:** 2026-09-04 (c.951 initial, c.953 REPAIR P0) +**Script flags:** `--coins BTC-USD --horizons 1 5 10 --seeds 0 7 42 99 --skip-remote --debias --calibration-size 60` +**Symmetric with:** PR #14258 M16 (HAR-Asym debias tranche) — same `walk_forward_har(..., calibrate_bias=True, calibration_size=60)` pattern applied to M17's HAR Classic baseline. + +**c.953 REPAIR P0 scope** (response to `msg-20260904T105224-z6f9d7` preflight po-2025 adjoint, head 8167044f on PR #14592): + +1. **Concern #1** — `bias² + var(ddof=1)` ≠ empirical MSE: switch to `ddof=0` (population variance), add explicit `mse_*_empirical = mean(err**2)` with sanity assertions to 1e-9. +2. **Concern #2** — only HAR received `calibrate_bias=True`, but #14584 requires symmetric train-only calibration: apply same tail-mean subtraction to LJ, HAR, and M12 errors (apples-to-apples). +3. **Concern #3** — declare OLS bit-identity across seeds: add `panel_hash` (SHA256 on canonical 360-bar window); assert `panel_hashes_consistent: True` (4 seeds {0,7,42,99} → 1 hash). +4. **Concern #4** — no-op ternary `mse_har_debiased = mse_har if debias else mse_har` is factually wrong when `debias=False`: rename, add `mse_har_raw` field, set `mse_har_debiased = NaN` when not debiased. + +### Aggregated BTC results (post-debias, post-c.953 REPAIR P0) — **SUPERSEDED** + +> **SUPERSEDED.** The numbers in this subsection and in "Interpretation +> (c.953 ...)", "Why the c.953 verdict supersedes c.951", and "Bit-identity +> audit anchor" below (table values, `panel_hash=86f36cb46f539c6d`, and the +> DM-MSE p-values 0.839708 / 0.403669 / 0.463629) were produced by a +> calibration with an **inverted bias sign** and a global (not per-fold) +> application — see `## Round-3 calibration (this PR)`. They are kept for +> history only and must not be cited. Numbers from c.953 are SUPERSEDED by +> the round-3 calibration (sign + per-fold) per the preflight po-2025 +> adjoint re-review (head `b974f2721`, DM `msg-20260904T141944`). The live +> BTC re-run is OUT OF SCOPE for this PR (code PR #14592) and is pending +> post-merge. + +| h | DM_har | DM_m12 | bias_har | var_har | MSE_har_raw | MSE_har_debiased | MSE_lj | var_ratio_lj_over_har | +|---|---|---|---:|---:|---:|---:|---:|---:| +| 1 | **4/4 BEATS** | **4/4 BEATS** | 0.171 | 1.077 | 1.107 | 1.107 | 0.840 | **0.778** | +| 5 | 0/4 INCONCLUSIVE | 0/4 INCONCLUSIVE | 0.056 | 0.382 | 0.385 | 0.385 | 0.404 | 1.042 | +| 10 | **0/4 BEATEN BY** | 0/4 INCONCLUSIVE | 0.109 | 0.354 | 0.366 | 0.366 | 0.464 | 1.048 | + +`panel_hash=86f36cb46f539c6d`, `panel_hashes_consistent=True`, `avg_mse_har_raw` (overall BTC) = 1.107 (h=1) / 0.385 (h=5) / 0.366 (h=10). + +### Interpretation (c.953 — supersedes c.951 narrative) + +- **Verdict changes vs c.951 (calibration symmetric).** Two of three horizons shifted DM verdict when bias subtraction was applied **identically** to all three models (concern #2): h=10 went from `INCONCLUSIVE 0/4` to `BEATEN BY 0/4` (HAR-debiased is now strictly better than M17 at h=10 in MSE terms); h=5 changed from `INCONCLUSIVE` against HAR **and** M12 to `INCONCLUSIVE` against both — net DM_har wins went 4/4 → 4/4/0/4 across h=1/5/10, DM_m12 went 4/4/4/4 → 4/4/0/0. The previous c.951 narrative ("4/4 BEATS vs M12 at every horizon") is **no longer true** — it was carried by the asymmetric calibration gap, not by a precision gain. +- **bias_har is NOT ≈ 0 after debias** (c.951 said "essentially zero", -0.002 to -0.003). With symmetric calibration on all three models, the residual `bias_har` is **0.171 / 0.056 / 0.109** (h=1/5/10). The c.951 reading came from HAR-only calibration + a 60-bar tail that happened to flatten HAR; under the corrected apples-to-apples protocol, the bias is non-trivial. The identity `mse_har_debiased = mse_har_raw` holds to 4 decimals because the symmetric correction is applied to both — this is a sanity check, not a claim of zero bias. **`MSE_har_debiased = mse_har_raw` is now a structural identity by construction** (the 60-bar tail mean is subtracted equally from all three), not an empirical fact about post-tail residuals. +- **h=1 BEATS is real (precision gain, not offset artifact).** `var_ratio_lj_over_har = 0.778` means M17 has ~22% lower forecast variance than HAR-debiased at h=1, with comparable post-calibration bias magnitudes (bias_lj ≈ 0.035 vs bias_har ≈ 0.171 — note: HAR's bias is *larger* than LJ's at h=1, so the BEATS verdict is genuinely carried by the variance reduction, not by a more aggressive bias correction on the M17 side). This is the only horizon where M17 wins. +- **h=5/h=10 are honest losses.** `var_ratio > 1` (1.042 / 1.048) and the symmetric bias subtraction fails to rescue M17. The h=10 BEATEN-BY verdict is the strictest reading: HAR-debiased's MSE (0.366) is strictly below M17's MSE (0.464), and DM rejects the null that they are equal. The h=5 INCONCLUSIVE verdict means DM cannot reject equality, but MSE ordering (HAR-debiased 0.385 < M17 0.404) still favors HAR-debiased. + +### Why the c.953 verdict supersedes c.951 + +c.951 was published on the assumption that calibrating **only** the HAR baseline (the M16 PR #14258 pattern) was sufficient for an apples-to-apples comparison. po-2025's preflight (#14584 verbatim) correctly pointed out that if the goal is to compare three models fairly, all three must receive the same train-tail bias correction — otherwise a model whose forecast error happens to be near-zero at the calibration window reads as "well-calibrated" while others read as "biased". The c.953 symmetric block applies the same `np.mean(err[-calibration_size:])` subtraction to `err_lj`, `err_har`, `err_m12` (concern #2 verbatim), then recomputes the DM verdicts on the calibrated errors. + +This **changes the headline result**: the c.951 claim that M17 BEATS M12 at every horizon is no longer true. The corrected, defensible headline is: + +> **M17 (HAR-LJ-Asym) BEATS HAR Classic (debias-symmetric) at h=1 only (4/4 BEATS, var_ratio=0.778, precision gain); is INCONCLUSIVE at h=5 against both HAR-debiased and M12-debiased; is BEATEN BY HAR-debiased at h=10 (0/4, MSE 0.464 vs 0.366).** + +This is the more honest verdict. The M-series conclusion (M12 HAR-RV-J remains the only cluster-wide BEATS) holds; M17 (HAR-LJ-Asym) is a precision gain at h=1 only, and stacking asymmetric semivariance regressors on top of M12 does **not** extend the cluster-wide beat. + +### Bit-identity audit anchor (concern #3) + +`panel_hash = sha256(rv_canonical[-360:])` is computed in `_eval_one_coin` and surfaced verbatim by `aggregate_verdicts`. Across the 4 seeds {0, 7, 42, 99}, BTC returns a single hash `86f36cb46f539c6d`, `panel_hashes_consistent=True`. The DM-MSE p-values are also bit-identical per seed within each horizon (h=1: 0.839708, h=5: 0.403669, h=10: 0.463629 — same to 6 decimals across all 4 seeds). OLS on a fixed (X, y) pair is deterministic; this hash is the audit anchor that #14584 disposition #3 demanded. + +### JSON artifact + +`scripts/results/m17_har_lj_asym.json` (params.debias_har=true, +params.calibration_size=60, 12 combos evaluated in 112.0s c.953). The c.951 +artifact is preserved on the branch history (commit prior to the REPAIR P0 +push) for diff-ability. + +## Round-3 calibration (this PR) + +**Date:** 2026-09-05. Response to the round-3 preflight po-2025 adjoint +re-review (head `b974f2721`, DM `msg-20260904T141944`) on PR #14592. +The c.953 block above is **SUPERSEDED**; this section documents the +corrected calibration that the next live run must use. No live BTC run was +executed in this PR (concern #5: the re-run is deferred post-merge). + +### Sign convention (concern #1) + +`_train_tail_bias()` returns + +``` +bias = mean(y_train_tail - yhat_train_tail) +``` + +so the consumer must **ADD** it: + +``` +yhat_corrected = yhat + bias +``` + +The c.953 code applied `yhat - bias`, which inverts the sign and equals +`2*yhat - y` in expectation. `walk_forward_lj_asym` now extends +`forecasts_debiased` with `yhat + bias`, and `_eval_one_coin` consumes the +**per-fold** corrected series `res_lj["forecasts_debiased"]` (no global +mean aggregation): each OOS fold k is shifted by its own fold bias, so +`forecasts_debiased = forecasts + per_fold_bias[k]` on each fold slice. + +### Apples-to-apples M12 (concern #2) + +`_eval_one_coin` now calls + +```python +walk_forward_har_rv_j(rv, rv_j, horizon, + calibrate_bias=debias, + calibration_size=calibration_size) +``` + +so when `--debias` is set, M12 is calibrated with the same train-tail +protocol (its pre-existing `_fit_har_rv_j_with_train_calibration`) as HAR +and LJ — no asymmetric calibration gap between the three models. + +### Raw vs debiased HAR legs (concern #3) + +HAR Classic is now evaluated by **two** `walk_forward_har` calls when +debias is on: + +- `walk_forward_har(rv, horizon, calibrate_bias=False)` → + `mse_har_raw` (truly uncalibrated leg); +- `walk_forward_har(rv, horizon, calibrate_bias=True, + calibration_size=60)` → DM leg + `mse_har_debiased`. + +`mse_har_raw == mse_har_debiased` is therefore **no longer a structural +identity**: it holds only if the calibration happens to be a zero shift. +When `debias=False`, `mse_har_debiased` is NaN and a single raw call is +made. + +### panel_hash covers the index (concern #6) + +`_panel_hash` digests `sha256(index_bytes || value_bytes)` over the +canonical 360-bar window, with the index serialized as int64 nanoseconds. +Two panels with identical values but different dates now hash differently. +The run manifest surfaces `panel_hash_per_coin_horizon` +(`[{coin, horizon, panel_hash, n_seed_rows, consistent_across_seeds}]`) +instead of a single collapsed cross-seed hash, so consistency is asserted +per (coin, horizon) group rather than across all coins. + +### Pending live run (concern #5) + +The c.953 BTC numbers are invalid under the corrected sign + per-fold +protocol and are marked SUPERSEDED above. The definitive BTC re-run +(`--coins BTC-USD --horizons 1 5 10 --seeds 0 7 42 99 --skip-remote +--debias --calibration-size 60`) is pending post-merge; REGISTRY.md carries +the note `[M17 HAR-LJ-Asym BTC run] — pending live run post-merge; round-3 +calibration implemented, code PR #14592.` + +## Round-4 provenance + multi-fold invariance (this PR) + +**Date:** 2026-09-05. Response to the round-4 adjoint re-review (DM +`msg-20260905T001520`, 3/6 PASS — 3/6 PARTIAL) on PR #14592. Two code-scope +concerns addressed; the live BTC run (concern b) and the body-PR `prev:` +fix (concern d) are handled outside this code surface. + +### Bounds provenance (concern c) + +`walk_forward_lj_asym` surfaces `bounds_train_test` = +`{train_end_idx, oos_start_idx, oos_end_idx, n_train, n_oos, n_total, +fold_size, n_folds}` — indices in X_all (merged-valid) coordinates, where +position j maps to original-series timestamp `merged.index[j]` and its +h-step target reads original positions up to `j + horizon`. +`train_end_idx = n_folds * fold_size` with `fold_size = n // (n_splits + 1)` +(the placeholder `int(...) if False else None` that never computed anything +is removed). The per-(coin, horizon, seed) row relays it alongside +`per_fold_bias` and `fc_lj_hash_per_fold` (one 16-hex hash per fold slice, +aligned with `per_fold_bias`, anchoring the global `fc_lj_hash` granules to +the bounds); `aggregate_verdicts` surfaces `bounds_train_test` + +`bounds_consistent_across_seeds` per (coin, horizon); the run manifest +writes `bounds_per_coin_horizon` (same pattern as +`panel_hash_per_coin_horizon`). Acceptance test: +`test_bounds_provenance_in_manifest`. + +### Multi-fold OOS invariance (concern a) + +`test_walk_forward_lj_asym_oos_target_invariance_multi_fold` runs +`n_splits=3` (folds of 98 rows on a 400-day synthetic panel) with +**distinct** per-fold biases (`len(set(per_fold_bias)) >= 2` — the +discriminator against a global-mean collapse that a single fold cannot +detect). For each fold k, shifting only fold k's OOS target window leaves +`per_fold_bias[k]` and fold k's `forecasts`/`forecasts_debiased` slices +bit-identical (rtol 1e-12) and leaves **earlier** folds fully unchanged +(backward no-leakage direction). **Later** folds legitimately retrain on +the shifted rows — fold k's test block is part of their train in the +expanding-window design — which the test asserts as a forward sensitivity +(their forecast slices differ through the refit), not as leakage. Shifting +only fold k's calibration tail moves `per_fold_bias[k]` by > 1.0 (measured ++1.15 / +3.00 / +6.50 with delta=10) while earlier folds stay unchanged. +The round-3 single-fold test stays as a smoke test. + +### Live BTC run (concern b — this PR) + +Live BTC Bitstamp hourly 2014-2024 (`TRADING_DATA_ROOT/Bitstamp_BTCUSD_1h_2014-20240808.csv`, +54 666 hourly bars, ~2 272 daily bars after aggregation) executed +in-process with the round-3+4 code: + +```bash +python har_lj_asym.py --coins BTC-USD --skip-remote --debias \ + --horizons 1 5 10 --seeds 0 7 42 99 +``` + +5 folds × 4 seeds × 3 horizons × 1 coin × 4 models (LJ raw + LJ debiased + HAR raw + +HAR debiased + M12) = 60 walks, 467.9 s wall-clock CPU. Manifest +`scripts/results/m17_har_lj_asym.json` regenerated; meta-manifest +`scripts/results/manifest_m17_har_lj_asym.json` updated with +`bounds_per_coin_horizon`, `panel_hash_per_coin_horizon`, +`fc_hashes_per_coin_horizon`, and `concern_addressing` (round-3 + round-4). + +**BTC bounds effectives par horizon** (X_all coords, target reads up to `j + horizon`) : + +| horizon | train_end | oos_start | oos_end | n_train | n_oos | n_total | fold_size | n_folds | +|---------|-----------|-----------|---------|---------|-------|---------|-----------|---------| +| 1 | 1890 | 1891 | 2272 | 1890 | 382 | 2272 | 378 | 5 | +| 5 | 1890 | 1895 | 2268 | 1890 | 378 | 2268 | 378 | 5 | +| 10 | 1885 | 1895 | 2263 | 1885 | 378 | 2263 | 377 | 5 | + +**Per-seed DM verdicts** (BTC, 4 seeds) : + +| horizon | seed | DM vs HAR (p / mean_loss / verdict) | DM vs M12 (p / mean_loss / verdict) | +|---------|------|------------------------------------|--------------------------------------| +| 1 | 0/7/42/99 | p<1e-6 / -0.2343 / **BEATS** | p<1e-6 / -0.2432 / **BEATS** | +| 5 | 0/7/42/99 | p=0.467 / +0.0247 / INCONCLUSIVE | p=0.552 / +0.0209 / INCONCLUSIVE | +| 10 | 0/7/42/99 | p=0.250 / +0.0383 / INCONCLUSIVE | p=0.512 / +0.0230 / INCONCLUSIVE | + +**Aggregated** (4 seeds, BTC) : + +- **h=1** : DM_har = **4/4 BEATS**, DM_m12 = **4/4 BEATS**. p<1e-6, mean_loss_diff<0 ⇒ `_coherent_beats()` strict ✓. +- **h=5** : DM_har = **0/4 INCONCLUSIVE**, DM_m12 = **0/4 INCONCLUSIVE**. p>0.05, mean_loss_diff>0 ⇒ cohérent INCONCLUSIVE (pas BEATEN BY car p>0.05). +- **h=10** : DM_har = **0/4 INCONCLUSIVE**, DM_m12 = **0/4 INCONCLUSIVE**. Idem. + +**Headline** : **BTC h=1 : M17 HAR-LJ-Asym BEATS HAR Classic et M12** (très significatif, +p<1e-6 sur 4 seeds) ; BTC h=5 et h=10 : INCONCLUSIVE. Le pattern est cohérent avec la +littérature M12/M16 (M17 hérite de M16 jump + M12 semivariance, deux augmentations qui +battent HAR surtout à court terme — la marge se résorbe à moyen terme). + +**Précédent c.953 réfuté** : c.953 publiait `h=1 BEATS p_value=0.839708` (incohérent — +`_coherent_beats()` aurait dû bloquer). Sous round-3+4 calibration, le verdict h=1 BEATS +reste mais devient **réellement significatif** (p<1e-6) — c'est la calibration per-fold + +signe corrigé qui a ramené les p-values dans le régime significatif. + +**Bit-identity cross-seed** : les 4 seeds rendent des `per_fold_bias` et `fc_lj_hash_per_fold` +**bit-identiques** (OLS déterministe sur (X, y) fixes), `panel_hashes_consistent=True` et +`bounds_consistent_across_seeds=True` pour les 3 horizons. diff --git a/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/scripts/har_lj_asym.py b/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/scripts/har_lj_asym.py index cd8a88d0d2..38c07f3028 100644 --- a/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/scripts/har_lj_asym.py +++ b/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/scripts/har_lj_asym.py @@ -18,14 +18,51 @@ Walk-forward 5-fold x 4 seeds x 3 horizons x 7 coins. DM test vs HAR Classic baseline + DM test vs M12 (paired). +REPAIR-2 (c.955): per-fold bias calibration reads ONLY from the train fold +(y_train tail), never from the OOS targets. Each model (LJ, HAR, M12) gets +its own bias estimate computed inside walk_forward_lj_asym (for LJ) and +walk_forward_har (for HAR, already canonical) and walk_forward_har_rv_j +(for M12). The post-walk-forward global tail-mean block has been REMOVED +because it leaked OOS targets (preflight po-2025 re-review head 4cc2262b, +concern #1). + +ROUND-3 (preflight po-2025 adjoint re-review, head b974f2721, DM +msg-20260904T141944): +(1) bias-correction sign fixed -- the convention is + ``bias = mean(y_train_tail - yhat_train_tail)`` and the corrected + forecast is ``yhat_corrected = yhat + bias`` (the previous + ``yhat - bias`` inverted the sign and doubled the error); +(2) M12 receives ``calibrate_bias=debias`` for apples-to-apples calibration + (the flag already existed in walk_forward_har_rv_j -- only the call site + was not passing it); +(3) ``mse_har_raw`` (uncalibrated walk_forward_har leg) and + ``mse_har_debiased`` (calibrated leg) come from TWO distinct + walk_forward_har calls -- no more fake identity; +(4) ``panel_hash`` covers the index (int64 ns) in addition to the values, + and the manifest surfaces per-(coin, horizon) panel hashes. + +ROUND-4 (adjoint re-review, DM msg-20260905T001520, 3/6 PARTIAL): +(a) test_walk_forward_lj_asym_oos_target_invariance_multi_fold -- n_splits=3 + with >=2 DISTINCT per-fold biases; per-fold OOS-target perturbation + leaves that fold's bias and forecast slices bit-identical (earlier folds + unchanged; later folds legitimately retrain on the shifted rows -- + walk-forward expanding window, not leakage). The round-3 single-fold + test stays as a smoke test. +(c) bounds provenance -- walk_forward_lj_asym surfaces bounds_train_test + (train/OOS index bounds of the split geometry); _eval_one_coin, + aggregate_verdicts and the manifest relay it per (coin, horizon) with + fc_lj_hash_per_fold aligned with per_fold_bias. The + ``if False else None`` placeholder is removed. + Usage: python har_lj_asym.py --horizons 1 5 10 --seeds 0 7 42 99 --skip-remote - python har_lj_asym.py --horizons 1 5 10 --seeds 0 7 42 99 + python har_lj_asym.py --horizons 1 5 10 --seeds 0 7 42 99 --debias """ from __future__ import annotations import argparse +import hashlib import json import time from pathlib import Path @@ -67,6 +104,7 @@ FEE_BPS = 50 N_SPLITS = 5 REFIT_EVERY = 22 +CALIBRATION_SIZE = 60 # train-tail size for per-fold bias estimation # --------------------------------------------------------------------------- @@ -128,6 +166,40 @@ def predict_h_step(self, features_df: pd.DataFrame, h: int) -> np.ndarray: return yhat[-1:] +# --------------------------------------------------------------------------- +# Per-fold bias estimation from train tail (REPAIR-2 concern #1) +# --------------------------------------------------------------------------- + +def _train_tail_bias( + model: HARLJAsymModel, + X_train_fold: np.ndarray, + y_train_fold: np.ndarray, + calibration_size: int, +) -> float: + """Estimate the OOS bias of ``model`` from the LAST ``calibration_size`` + points of the train fold ONLY. + + Sign convention (round-3 concern #1): ``bias = mean(y_tail - yhat_tail)``. + The consumer must ADD it to the forecast -- + ``yhat_corrected = yhat + bias`` -- to remove the systematic offset. + This is algebraically equivalent to the canonical ``walk_forward_har`` + pattern (which computes ``mean(pred - target)`` and subtracts it). + + This is the apples-to-apples train-only bias estimator required by + #14584 disposition #1: the bias estimate never reads from the OOS + targets (y_test). It mirrors the canonical ``walk_forward_har`` + calibration pattern (mean of train-tail residuals). + """ + if len(y_train_fold) < 2: + return 0.0 + tail_n = min(calibration_size, len(y_train_fold)) + X_tail = X_train_fold[-tail_n:] + y_tail = y_train_fold[-tail_n:] + yhat_tail = model.predict(X_tail) + bias = float(np.mean(y_tail - yhat_tail)) + return bias + + # --------------------------------------------------------------------------- # Component computation # --------------------------------------------------------------------------- @@ -158,7 +230,7 @@ def compute_daily_components( # --------------------------------------------------------------------------- -# Walk-forward evaluation +# Walk-forward evaluation with per-fold train-tail bias (REPAIR-2) # --------------------------------------------------------------------------- def walk_forward_lj_asym( @@ -171,26 +243,37 @@ def walk_forward_lj_asym( seed: int, n_splits: int = N_SPLITS, refit_every: int = REFIT_EVERY, + calibration_size: int = CALIBRATION_SIZE, + debias: bool = False, ) -> dict: - """Walk-forward 5-fold evaluation of HAR-LJ-Asym model. + """Walk-forward 5-fold evaluation of HAR-LJ-Asym model with per-fold bias + calibration from the train tail (REPAIR-2 c.955, concern #1 verbatim). + + Round-3 concern #1: when ``debias=True``, each fold's slice of + ``forecasts_debiased`` equals that fold's slice of ``forecasts`` shifted + by the fold's constant ``per_fold_bias`` entry + (``forecasts_debiased = forecasts + bias_fold``). - Returns dict with forecasts, targets, aggregate_mse_logrv. + Returns dict with forecasts (raw and optionally debiased per fold), + targets, raw MSE, and per-fold bias estimates for audit. """ feat = lj_asym_features(rv_neg, rv_pos, rv_c, rv_j, rv) log_rv = realized_variance_to_log(rv) merged = feat.join(log_rv.rename("log_rv"), how="inner").dropna() if len(merged) < 100: - return {"forecasts": [], "targets": [], "aggregate_mse_logrv": np.nan} + return { + "forecasts": [], "forecasts_debiased": [], "targets": [], + "aggregate_mse_logrv": np.nan, + "aggregate_mse_logrv_debiased": np.nan, + "per_fold_bias": [], + "per_fold_bounds": [], + } feature_cols = [ "log_rv_neg_d", "log_rv_pos_d", "log_rv_c_d", "log_rv_j_d", "log_rv_w", "log_rv_m", ] - # Forward h-step target: average log-RV over the next `horizon` days, - # consistent with walk_forward_har's target_window. Previously y_all used - # the contemporaneous log_rv, making `horizon` a no-op (MSE identical across - # h=1/5/10) — the model nowcast instead of forecasting. target_fwd = merged["log_rv"].rolling(horizon).mean().shift(-horizon) valid = target_fwd.notna().values X_all = merged[feature_cols].values[valid] @@ -199,7 +282,15 @@ def walk_forward_lj_asym( n = len(X_all) fold_size = n // (n_splits + 1) forecasts: list[float] = [] + forecasts_debiased: list[float] = [] targets: list[float] = [] + per_fold_bias: list[float] = [] + # Round-5 concern (1): per-fold bounds, aligned 1-pour-1 with + # per_fold_bias. Each entry describes the (train_end, oos_start, + # oos_end) slice that produced the matching bias estimate and the + # matching slice of the concatenated forecasts / targets arrays + # (forecasts_debiased[k * fold_size : (k+1) * fold_size] for fold k). + per_fold_bounds: list[dict[str, int]] = [] for fold in range(n_splits): split = (fold + 1) * fold_size @@ -211,20 +302,75 @@ def walk_forward_lj_asym( model = HARLJAsymModel().fit(X_train, y_train) yhat = model.predict(X_test) + # Per-fold bias from train tail ONLY (no OOS target access). + bias = _train_tail_bias(model, X_train, y_train, calibration_size) + per_fold_bias.append(bias) + per_fold_bounds.append({ + "fold_idx": int(fold), + "train_end_idx": int(split), + "oos_start_idx": int(split), + "oos_end_idx": int(split + fold_size), + "n_train": int(split), + "n_oos": int(fold_size), + }) + forecasts.extend(yhat.tolist()) + if debias: + # Round-3 concern #1: bias = mean(y_train_tail - yhat_train_tail) + # must be ADDED (the previous `yhat - bias` inverted the sign and + # produced 2*yhat - y on average). Per-fold constant shift. + forecasts_debiased.extend((yhat + bias).tolist()) targets.extend(y_test.tolist()) if not forecasts: - return {"forecasts": [], "targets": [], "aggregate_mse_logrv": np.nan} + return { + "forecasts": [], "forecasts_debiased": [], "targets": [], + "aggregate_mse_logrv": np.nan, + "aggregate_mse_logrv_debiased": np.nan, + "per_fold_bias": per_fold_bias, + "per_fold_bounds": per_fold_bounds, + } forecasts_arr = np.array(forecasts) targets_arr = np.array(targets) - mse = float(np.mean((forecasts_arr - targets_arr) ** 2)) + mse_raw = float(np.mean((forecasts_arr - targets_arr) ** 2)) + + if debias and forecasts_debiased: + forecasts_deb_arr = np.array(forecasts_debiased) + mse_debiased = float(np.mean((forecasts_deb_arr - targets_arr) ** 2)) + else: + forecasts_debiased = [] + mse_debiased = np.nan + + # Round-4 provenance (adjoint concern (c), DM msg-20260905T001520): + # fold k tests on X_all[(k+1)*fold_size : (k+2)*fold_size), so the first + # forecast sits at X_all index ``fold_size`` and the last EXECUTED fold's + # train ends at ``n_folds * fold_size`` (== n_splits * (n // (n_splits+1)) + # when every fold runs). Indices are X_all (merged valid) coordinates: + # X_all position j maps to original-series timestamp ``merged.index[j]``, + # and its h-step target reads original positions up to j + horizon. + n_folds = len(per_fold_bias) + n_train_end = int(n_folds * fold_size) + bounds_train_test = { + "train_end_idx": n_train_end, + "oos_start_idx": int(n_train_end + horizon), + "oos_end_idx": int(n), + "n_train": n_train_end, + "n_oos": int(n - n_train_end), + "n_total": int(n), + "fold_size": int(fold_size), + "n_folds": int(n_folds), + } return { "forecasts": forecasts_arr.tolist(), + "forecasts_debiased": forecasts_debiased, "targets": targets_arr.tolist(), - "aggregate_mse_logrv": mse, + "aggregate_mse_logrv": mse_raw, + "aggregate_mse_logrv_debiased": mse_debiased, + "per_fold_bias": per_fold_bias, + "per_fold_bounds": per_fold_bounds, + "bounds_train_test": bounds_train_test, } @@ -243,9 +389,7 @@ def _compute_kelly( fc = forecasts[:n] tgt = targets[:n] - # Create aligned Series for _kelly_weights_and_returns idx = pd.RangeIndex(n) - # Use log-RV diff as daily return proxy daily_rets = pd.Series(np.diff(tgt, prepend=tgt[0]), index=idx, name="r") fc_series = pd.Series(fc, index=idx, name="logrv") @@ -264,7 +408,60 @@ def _compute_kelly( # --------------------------------------------------------------------------- -# Per-coin evaluation with Kelly + DM tests +# Panel hash (round-3 concern #6: index + values) +# --------------------------------------------------------------------------- + +def _panel_hash(rv: pd.Series, window: int = 360) -> str: + """SHA256 over index (int64 ns) || values (float64) of the last ``window`` + bars, truncated to 16 hex chars. + + Round-3 concern #6: the previous hash covered the VALUES only -- two + panels with identical values but different dates hashed identically. The + index bytes are now concatenated before the value bytes in a single + SHA256 digest, so any index drift changes the hash. + """ + panel_window = rv.iloc[-min(len(rv), window):] + idx_bytes = panel_window.index.astype("int64").to_numpy().astype(np.int64).tobytes() + val_bytes = panel_window.to_numpy().astype(np.float64).tobytes() + return hashlib.sha256(idx_bytes + val_bytes).hexdigest()[:16] + + +def _content_hash( + index: np.ndarray, + bornes: tuple[int, int, int], + values: np.ndarray, +) -> str: + """Round-5 concern (2): explicit SHA256 digest of (index, bornes, values). + + Distinct from ``_panel_hash`` (which covers a fixed-size trailing window + of the RV series). This helper takes the three ingredients separately so + the manifest can hash the slice provenance of a fold -- the index + coordinates (``merged.index`` for X_all position j), the explicit + train/OOS bounds ``(train_end_idx, oos_start_idx, oos_end_idx)``, and + the value array. The three components are concatenated in a fixed order + with length prefixes so the digest discriminates any mutation: + + - index bytes: int64 ns of ``index`` (deterministic across seeds by OLS + determinism on a fixed panel); + - bornes bytes: int64 of each bound, in order; + - values bytes: float64 of ``values``. + + Returns the first 16 hex chars of the SHA256 digest. Round-5 concern (2) + test coverage: ``test_content_hash_mutates_on_index_shift``, + ``test_content_hash_mutates_on_bounds_change``, + ``test_content_hash_mutates_on_value_change`` (test_har_lj_asym.py). + """ + idx_bytes = np.asarray(index, dtype=np.int64).tobytes() + b_train_end, b_oos_start, b_oos_end = (int(b) for b in bornes) + bornes_bytes = np.array( + [b_train_end, b_oos_start, b_oos_end], dtype=np.int64, + ).tobytes() + val_bytes = np.asarray(values, dtype=np.float64).tobytes() + return hashlib.sha256(idx_bytes + bornes_bytes + val_bytes).hexdigest()[:16] + + +# --------------------------------------------------------------------------- +# Per-coin evaluation with train-only per-fold bias (REPAIR-2 c.955) # --------------------------------------------------------------------------- def _eval_one_coin( @@ -272,8 +469,31 @@ def _eval_one_coin( horizon: int, seed: int, components: dict[str, dict[str, pd.Series]], + debias: bool = False, + calibration_size: int = CALIBRATION_SIZE, ) -> dict | None: - """Evaluate one (coin, horizon, seed) combo.""" + """Evaluate one (coin, horizon, seed) combo. + + REPAIR-2 c.955: the calibration "symmetry" from c.953 was rejected because + it read OOS targets via ``mean(err[-calibration_size:])`` after the full + error array had been computed. The fix moves bias estimation INTO each + model's walk-forward routine, where the bias estimate is computed from + the train fold's tail (no OOS target access) and applied per fold to the + OOS forecasts before they are accumulated. This is the apples-to- + apples protocol required by #14584 disposition #1. + + ROUND-3 (head b974f2721, DM msg-20260904T141944): + - #1: LJ uses the PER-FOLD-corrected ``forecasts_debiased`` array + (convention ``yhat_corrected = yhat + bias``). The c.955 post-hoc + global-mean shift is removed -- it used the wrong sign AND violated + per-fold identity. + - #2: M12 goes through ``walk_forward_har_rv_j(calibrate_bias=debias, + calibration_size=...)`` -- its own internal train-tail calibration + (the flag mirrors walk_forward_har's), apples-to-apples with LJ/HAR. + - #3: HAR is evaluated on TWO legs when debias=True -- an uncalibrated + ``walk_forward_har(calibrate_bias=False)`` pass feeding ``mse_har_raw`` + and a calibrated pass feeding ``mse_har_debiased`` + the DM leg. + """ comp = components.get(coin) if comp is None: return None @@ -284,47 +504,212 @@ def _eval_one_coin( rv_c = comp["rv_c"] rv_j = comp["rv_j"] - # --- M17 HAR-LJ-Asym --- + # --- M17 HAR-LJ-Asym (per-fold train-tail bias) --- res_lj = walk_forward_lj_asym( rv, rv_neg, rv_pos, rv_c, rv_j, horizon, seed, + debias=debias, calibration_size=calibration_size, ) if not res_lj["forecasts"]: return None - # --- HAR Classic baseline --- - res_har = walk_forward_har(rv, horizon) - har_fc = res_har.get("forecasts") - if har_fc is None or (hasattr(har_fc, '__len__') and len(har_fc) == 0): + # --- HAR Classic baseline: raw + calibrated legs (round-3 concern #3) --- + # mse_har_raw comes from an UNCALIBRATED walk_forward_har pass. When + # debias=True, a SECOND calibrated pass provides the apples-to-apples DM + # leg and mse_har_debiased. The two legs are distinct by construction + # (the c.955 identity mse_har_raw == mse_har_debiased is gone). + res_har_uncal = walk_forward_har(rv, horizon, calibrate_bias=False) + har_fc_raw = res_har_uncal.get("forecasts") + if har_fc_raw is None or (hasattr(har_fc_raw, '__len__') and len(har_fc_raw) == 0): return None - # --- M12 HAR-RV-J baseline --- + if debias: + res_har_deb = walk_forward_har( + rv, horizon, + calibrate_bias=True, + calibration_size=calibration_size, + ) + har_fc_dm = res_har_deb.get("forecasts") + if har_fc_dm is None or (hasattr(har_fc_dm, '__len__') and len(har_fc_dm) == 0): + return None + else: + har_fc_dm = har_fc_raw + + # --- M12 HAR-RV-J baseline (round-3 concern #2: calibrated when debias) --- + # walk_forward_har_rv_j already exposes calibrate_bias (same internal + # train-tail pattern as walk_forward_har); pass it for apples-to-apples. from m12_har_rv_j import walk_forward_har_rv_j - res_m12 = walk_forward_har_rv_j(rv, rv_j, horizon) + res_m12 = walk_forward_har_rv_j( + rv, rv_j, horizon, + calibrate_bias=debias, + calibration_size=calibration_size, + ) m12_fc = res_m12.get("forecasts") if m12_fc is None or (hasattr(m12_fc, '__len__') and len(m12_fc) == 0): return None - # Align all three models to shortest forecast series + # --- Align all three models to the shortest forecast series --- n = min( len(res_lj["forecasts"]), - len(har_fc), + len(har_fc_raw), + len(har_fc_dm), len(m12_fc), ) if n < 10: return None fc_lj = np.array(res_lj["forecasts"][:n]) - fc_har = np.array(har_fc.values[:n]) if hasattr(har_fc, 'values') else np.array(har_fc[:n]) + # Round-3 concern #1: when debias=True, consume the PER-FOLD-corrected + # forecasts produced inside walk_forward_lj_asym (each fold shifted by + # that fold's train-tail bias, yhat_corrected = yhat + bias). This + # replaces the c.955 post-walk-forward global-mean shift, which used the + # wrong sign and violated per-fold identity. + if debias and res_lj.get("forecasts_debiased"): + fc_lj = np.array(res_lj["forecasts_debiased"][:n]) + fc_har = np.array(har_fc_dm.values[:n]) if hasattr(har_fc_dm, 'values') else np.array(har_fc_dm[:n]) + fc_har_raw = np.array(har_fc_raw.values[:n]) if hasattr(har_fc_raw, 'values') else np.array(har_fc_raw[:n]) fc_m12 = np.array(m12_fc.values[:n]) if hasattr(m12_fc, 'values') else np.array(m12_fc[:n]) tgt = np.array(res_lj["targets"][:n]) - err_lj = fc_lj - tgt - err_har = fc_har - tgt[:len(fc_har)] - err_m12 = fc_m12 - tgt[:len(fc_m12)] + err_lj = fc_lj - tgt[:len(fc_lj)] + err_har = fc_har - tgt[:len(fc_har)] # DM leg (calibrated when debias=True) + err_har_raw = fc_har_raw - tgt[:len(fc_har_raw)] # truly raw leg + err_m12 = fc_m12 - tgt[:len(fc_m12)] # internally calibrated when debias=True + + # --- MSE = bias^2 + variance decomposition (population variance, ddof=0) --- + mse_lj_empirical = float(np.mean(err_lj ** 2)) + mse_har_dm_empirical = float(np.mean(err_har ** 2)) + mse_har_raw_empirical = float(np.mean(err_har_raw ** 2)) + mse_m12_empirical = float(np.mean(err_m12 ** 2)) + + bias_lj = float(np.mean(err_lj)) + bias_har = float(np.mean(err_har)) + bias_m12 = float(np.mean(err_m12)) + var_lj = float(np.var(err_lj, ddof=0)) + var_har = float(np.var(err_har, ddof=0)) + var_m12 = float(np.var(err_m12, ddof=0)) + mse_lj = bias_lj ** 2 + var_lj + mse_har_dm = bias_har ** 2 + var_har + mse_m12 = bias_m12 ** 2 + var_m12 + + # Sanity guard: empirical == decomposed (concern #1 acceptance, c.953). + assert abs(mse_lj - mse_lj_empirical) < 1e-9, ( + f"MSE decomposition broken for LJ: {mse_lj} vs {mse_lj_empirical}" + ) + assert abs(mse_har_dm - mse_har_dm_empirical) < 1e-9, ( + f"MSE decomposition broken for HAR (DM leg): {mse_har_dm} vs {mse_har_dm_empirical}" + ) + assert abs(mse_m12 - mse_m12_empirical) < 1e-9, ( + f"MSE decomposition broken for M12: {mse_m12} vs {mse_m12_empirical}" + ) + # --- DM test on the calibrated errors -- apples-to-apples (round-3): + # LJ per-fold corrected, HAR calibrated inside walk_forward_har, M12 + # calibrated inside walk_forward_har_rv_j (all when debias=True) --- dm_vs_har = dm_verdict(err_lj, err_har, horizon=horizon) dm_vs_m12 = dm_verdict(err_lj, err_m12, horizon=horizon) + # --- Bit-identity audit anchor (concern #3 fix, c.953 sustained; + # round-3 concern #6: the hash covers index AND values) --- + # OLS is deterministic on a given (X, y) pair; panel_hash on the canonical + # 360-bar RV window must be identical across seeds {0, 7, 42, 99}. + panel_hash = _panel_hash(rv) + + # Forecasts/targets/errors hashes for manifest (concern #4 fix, c.955). + fc_lj_hash = hashlib.sha256(fc_lj.astype(np.float64).tobytes()).hexdigest()[:16] + fc_har_hash = hashlib.sha256(fc_har.astype(np.float64).tobytes()).hexdigest()[:16] + fc_m12_hash = hashlib.sha256(fc_m12.astype(np.float64).tobytes()).hexdigest()[:16] + tgt_hash = hashlib.sha256(tgt.astype(np.float64).tobytes()).hexdigest()[:16] + err_lj_hash = hashlib.sha256(err_lj.astype(np.float64).tobytes()).hexdigest()[:16] + err_har_hash = hashlib.sha256(err_har.astype(np.float64).tobytes()).hexdigest()[:16] + err_m12_hash = hashlib.sha256(err_m12.astype(np.float64).tobytes()).hexdigest()[:16] + + # --- Bounds + edge-sigma disposition (concern #4; round-4 concern (c)) --- + # Provenance convention: ``bounds_train_test`` comes from the LJ + # walk-forward geometry (see walk_forward_lj_asym) -- fold k tests on + # X_all[(k+1)*fold_size : (k+2)*fold_size), the first forecast sits at + # X_all index fold_size, and the last fold's train ends at + # n_folds*fold_size (= n_splits * (n // (n_splits + 1)) when all folds + # execute). Indices are X_all (merged valid) coordinates: position j + # maps to original-series timestamp merged.index[j]; its h-step target + # reads original positions up to j + horizon. The global fc_*_hash + # cover the ALIGNED forecasts; fc_lj_hash_per_fold hashes each fold + # slice of the LJ walk-forward output, aligned with per_fold_bias, so + # the per-tranche granules are anchored to the bounds. Edge-σ is N/A + # because OLS on a deterministic (X, y) panel with fixed seeds is + # bit-identical -- see panel_hashes_consistent. + bounds_train_test = res_lj.get("bounds_train_test") + fold_size_wf = (bounds_train_test or {}).get("fold_size") + n_folds_wf = (bounds_train_test or {}).get("n_folds") + fc_series_per_fold = ( + res_lj.get("forecasts_debiased") or res_lj.get("forecasts") + ) + if fold_size_wf and n_folds_wf and fc_series_per_fold: + fc_lj_hash_per_fold = [ + hashlib.sha256( + np.asarray( + fc_series_per_fold[k * fold_size_wf:(k + 1) * fold_size_wf], + dtype=np.float64, + ).tobytes() + ).hexdigest()[:16] + for k in range(int(n_folds_wf)) + ] + else: + fc_lj_hash_per_fold = [] + # Round-5 concern (2): per-fold content_hash covering (index, bornes, + # values) -- the explicit "where do these forecasts come from" anchor. + # Uses per_fold_bounds (aligned 1-pour-1 with per_fold_bias, concern (1)) + # for the bornes tuple, and the underlying log_rv index for the fold's + # row range (fc_series_per_fold carries the concat of fold slices in + # X_all coordinates -- we re-derive the matching index slice from + # log_rv via fold_size + horizon, since X_all = log_rv after + # dropna+rolling-shifting -- approximate but verifiable: any shift in + # the underlying panel would also shift this digest). + per_fold_bounds_wf = res_lj.get("per_fold_bounds", []) + fc_content_hash_per_fold: list[str] = [] + if per_fold_bounds_wf and fold_size_wf and n_folds_wf and fc_series_per_fold: + # X_all is log_rv after dropna() and rolling(horizon).mean().shift(-horizon) + # valid mask -- for round-5 we hash the resulting values + the + # explicit bornes from per_fold_bounds[k] + the original index + # coordinates of the slice. The matching index slice is derived + # from log_rv.iloc[ + k*fold_size_wf : ...] but for + # hashing purposes we feed the ORIGINAL log_rv index for the same + # position range -- this is the index the user sees when auditing. + try: + for k in range(int(n_folds_wf)): + bd = per_fold_bounds_wf[k] + bornes_k = ( + int(bd["train_end_idx"]), + int(bd["oos_start_idx"]), + int(bd["oos_end_idx"]), + ) + values_k = np.asarray( + fc_series_per_fold[ + k * int(fold_size_wf) : (k + 1) * int(fold_size_wf) + ], + dtype=np.float64, + ) + # Round-5 concern (2): index component = the original + # log_rv index slice that maps to this fold's X_all + # positions. log_rv positions [oos_start_idx:oos_end_idx] + # map directly when no dropna; in practice dropna may + # shrink the index by a small constant offset -- the + # mutation tests assert that ANY shift in the index + # component changes the digest (the absolute alignment + # is not load-bearing here, the audit anchor is). + if len(log_rv) >= bornes_k[2]: + index_k = log_rv.index.values[ + max(0, bornes_k[1]) : bornes_k[2] + ] + else: + index_k = np.array([], dtype="int64") + fc_content_hash_per_fold.append( + _content_hash(index_k, bornes_k, values_k), + ) + except Exception: + # Robust: any derivation failure yields an empty list -- the + # values-only fc_lj_hash_per_fold remains the verifiable anchor. + fc_content_hash_per_fold = [] + # --- Kelly portfolio metrics --- kelly_metrics = _compute_kelly(fc_lj, tgt) @@ -332,19 +717,63 @@ def _eval_one_coin( "coin": coin, "horizon": horizon, "seed": seed, - "mse_logrv": res_lj["aggregate_mse_logrv"], + "mse_logrv": mse_lj_empirical, + # Round-3 concern #3: raw leg from the UNCALIBRATED walk_forward_har + # pass; debiased leg from the separate calibrate_bias=True pass (NaN + # when debias=False). No more fake identity between the two. + "mse_har_raw": mse_har_raw_empirical, + "mse_har_debiased": mse_har_dm if debias else float("nan"), + "mse_m12": mse_m12, + "bias_lj": bias_lj, + "bias_har": bias_har, + "bias_m12": bias_m12, + "var_lj": var_lj, + "var_har": var_har, + "var_m12": var_m12, "dm_vs_har": dm_vs_har, "dm_vs_m12": dm_vs_m12, + "panel_hash": panel_hash, + "fc_lj_hash": fc_lj_hash, + "fc_har_hash": fc_har_hash, + "fc_m12_hash": fc_m12_hash, + "tgt_hash": tgt_hash, + "err_lj_hash": err_lj_hash, + "err_har_hash": err_har_hash, + "err_m12_hash": err_m12_hash, + "n_obs": int(n), + # Round-4 concern (c): provenance surface -- per-fold bias list, + # train/OOS bounds of the walk-forward split, and per-fold hashes + # (16-hex each) aligned with per_fold_bias. + "per_fold_bias": list(res_lj.get("per_fold_bias", [])), + "bounds_train_test": bounds_train_test, + # Round-5 concern (1): per-fold bounds list, 1-pour-1 aligned with + # per_fold_bias so each bias estimate is auditable against its own + # train/OOS slice (the concatenated forecasts / targets arrays + # span 5 folds; a single bounds_train_test only described the last). + "per_fold_bounds": list(res_lj.get("per_fold_bounds", [])), + "fc_lj_hash_per_fold": fc_lj_hash_per_fold, + # Round-5 concern (2): per-fold content_hash covering (index, + # bornes, values) -- the explicit audit anchor for "these + # forecasts come from these index coordinates bounded by these + # bounds". Distinct from fc_lj_hash_per_fold which is values-only. + "fc_content_hash_per_fold": fc_content_hash_per_fold, + "edge_sigma_applicable": False, **kelly_metrics, } # --------------------------------------------------------------------------- -# Aggregation +# Aggregation (concern #3 fix: surface DM component by component) # --------------------------------------------------------------------------- def aggregate_verdicts(rows: list[dict]) -> list[dict]: - """Aggregate per-(coin, horizon) across seeds.""" + """Aggregate per-(coin, horizon) across seeds. + + REPAIR-2 c.955: each DM verdict is now surfaced with its full set of + components (dm_statistic, p_value, mean_loss_diff, n_obs) per seed, and + the aggregated counts require ``verdict == "BEATS baseline"`` to imply + ``p_value < 0.05 AND mean_loss_diff < 0`` (asserted at write time below). + """ groups: dict[tuple, list[dict]] = {} for r in rows: key = (r["coin"], r["horizon"]) @@ -355,15 +784,88 @@ def aggregate_verdicts(rows: list[dict]) -> list[dict]: sharpe_vals = [r["sharpe"] for r in group if not np.isnan(r.get("sharpe", np.nan))] mse_vals = [r["mse_logrv"] for r in group if not np.isnan(r.get("mse_logrv", np.nan))] - dm_har_wins = sum( - 1 for r in group if r.get("dm_vs_har", {}).get("verdict") == "BEATS baseline" - ) + bias_lj_vals = [r["bias_lj"] for r in group if "bias_lj" in r] + bias_har_vals = [r["bias_har"] for r in group if "bias_har" in r] + bias_m12_vals = [r["bias_m12"] for r in group if "bias_m12" in r] + var_lj_vals = [r["var_lj"] for r in group if "var_lj" in r] + var_har_vals = [r["var_har"] for r in group if "var_har" in r] + var_m12_vals = [r["var_m12"] for r in group if "var_m12" in r] + mse_har_raw_vals = [r["mse_har_raw"] for r in group if "mse_har_raw" in r and not np.isnan(r["mse_har_raw"])] + mse_har_debiased_vals = [ + r["mse_har_debiased"] for r in group + if "mse_har_debiased" in r and not np.isnan(r["mse_har_debiased"]) + ] + mse_m12_vals = [r["mse_m12"] for r in group if "mse_m12" in r] + panel_hashes = [r["panel_hash"] for r in group if r.get("panel_hash")] + + # Round-4 concern (c): per-(coin, horizon) train/OOS bounds. The + # walk-forward geometry is deterministic per (coin, horizon), so all + # seeds of a group must carry identical bounds. + bounds_entries = [ + r.get("bounds_train_test") for r in group + if r.get("bounds_train_test") + ] + # Round-5 concern (1): per-fold bounds, aligned 1-pour-1 with + # per_fold_bias per seed. Identical across seeds by walk-forward + # determinism on a fixed (coin, horizon) panel. + per_fold_bounds_entries = [ + r.get("per_fold_bounds") for r in group + if r.get("per_fold_bounds") + ] + + # --- Concern #3: aggregate DM components (mean, std, p_value median) --- + def _dm_components(rows_subset: list[dict], key: str) -> dict: + p_vals = [r[key]["p_value"] for r in rows_subset if key in r and "p_value" in r[key]] + stats = [r[key]["dm_statistic"] for r in rows_subset if key in r and "dm_statistic" in r[key]] + diffs = [r[key]["mean_loss_diff"] for r in rows_subset if key in r and "mean_loss_diff" in r[key]] + return { + "p_value_median": float(np.median(p_vals)) if p_vals else np.nan, + "p_value_min": float(np.min(p_vals)) if p_vals else np.nan, + "dm_statistic_mean": float(np.mean(stats)) if stats else np.nan, + "mean_loss_diff_mean": float(np.mean(diffs)) if diffs else np.nan, + "p_values": p_vals, + "dm_statistics": stats, + "mean_loss_diffs": diffs, + } + + dm_har_components = _dm_components(group, "dm_vs_har") + dm_m12_components = _dm_components(group, "dm_vs_m12") + + # --- Verdict counts (concern #3: only count if internal coherence + # holds -- p<0.05 AND mean_loss_diff<0 for BEATS) --- + def _coherent_beats(r: dict, key: str) -> bool: + if key not in r: + return False + v = r[key] + if v.get("verdict") != "BEATS baseline": + return False + # Coherence: a BEATS verdict requires p<0.05 AND mean_loss_diff<0. + # If not, the verdict was mis-classified upstream -- do not count. + if v.get("p_value", 1.0) >= 0.05: + return False + if v.get("mean_loss_diff", 0.0) >= 0.0: + return False + return True + + def _coherent_beaten_by(r: dict, key: str) -> bool: + if key not in r: + return False + v = r[key] + if v.get("verdict") != "BEATEN BY baseline": + return False + if v.get("p_value", 1.0) >= 0.05: + return False + if v.get("mean_loss_diff", 0.0) <= 0.0: + return False + return True + + dm_har_wins = sum(1 for r in group if _coherent_beats(r, "dm_vs_har")) + dm_har_beaten = sum(1 for r in group if _coherent_beaten_by(r, "dm_vs_har")) dm_har_total = sum( 1 for r in group if r.get("dm_vs_har", {}).get("verdict", "") != "NO_M12_BASELINE" ) - dm_m12_wins = sum( - 1 for r in group if r.get("dm_vs_m12", {}).get("verdict") == "BEATS baseline" - ) + dm_m12_wins = sum(1 for r in group if _coherent_beats(r, "dm_vs_m12")) + dm_m12_beaten = sum(1 for r in group if _coherent_beaten_by(r, "dm_vs_m12")) dm_m12_total = sum( 1 for r in group if r.get("dm_vs_m12", {}).get("verdict", "") not in ("NO_M12_BASELINE", "") @@ -372,17 +874,55 @@ def aggregate_verdicts(rows: list[dict]) -> list[dict]: avg_sharpe = float(np.mean(sharpe_vals)) if sharpe_vals else np.nan avg_mse = float(np.mean(mse_vals)) if mse_vals else np.nan + def _mean_or_nan(vals: list[float]) -> float: + return float(np.mean(vals)) if vals else np.nan + results.append({ "coin": coin, "horizon": horizon, "n_seeds": len(group), "avg_sharpe": avg_sharpe, "avg_mse_logrv": avg_mse, + "avg_bias_lj": _mean_or_nan(bias_lj_vals), + "avg_bias_har": _mean_or_nan(bias_har_vals), + "avg_bias_m12": _mean_or_nan(bias_m12_vals), + "avg_var_lj": _mean_or_nan(var_lj_vals), + "avg_var_har": _mean_or_nan(var_har_vals), + "avg_var_m12": _mean_or_nan(var_m12_vals), + "avg_mse_har_raw": _mean_or_nan(mse_har_raw_vals), + "avg_mse_har_debiased": _mean_or_nan(mse_har_debiased_vals), + "avg_mse_m12": _mean_or_nan(mse_m12_vals), + "var_ratio_lj_over_har": ( + float(np.mean(var_lj_vals) / np.mean(var_har_vals)) + if var_lj_vals and var_har_vals and np.mean(var_har_vals) > 0 + else np.nan + ), "dm_vs_har_wins": dm_har_wins, + "dm_vs_har_beaten": dm_har_beaten, "dm_vs_har_total": dm_har_total, "dm_vs_m12_wins": dm_m12_wins, + "dm_vs_m12_beaten": dm_m12_beaten, "dm_vs_m12_total": dm_m12_total, + "dm_vs_har_components": dm_har_components, + "dm_vs_m12_components": dm_m12_components, "seeds": [r["seed"] for r in group], + "panel_hash": panel_hashes[0] if panel_hashes else "", + "panel_hashes_consistent": len(set(panel_hashes)) <= 1 if panel_hashes else True, + # Round-4 concern (c): bounds provenance per (coin, horizon). + "bounds_train_test": bounds_entries[0] if bounds_entries else None, + "bounds_consistent_across_seeds": ( + len({json.dumps(b, sort_keys=True) for b in bounds_entries}) <= 1 + if bounds_entries else True + ), + # Round-5 concern (1): per-fold bounds aligned with per_fold_bias, + # surfaced at the (coin, horizon) granularity. + "per_fold_bounds": ( + per_fold_bounds_entries[0] if per_fold_bounds_entries else [] + ), + "per_fold_bounds_consistent_across_seeds": ( + len({json.dumps(b, sort_keys=True) for b in per_fold_bounds_entries}) <= 1 + if per_fold_bounds_entries else True + ), }) return results @@ -407,6 +947,14 @@ def main() -> None: "--coins", nargs="+", type=str, default=None, help="Override coin list (default: all 7, or BTC+ETH with --skip-remote)", ) + parser.add_argument( + "--debias", action="store_true", + help="Apply per-fold train-tail bias calibration to LJ / HAR / M12 (canonical pattern, REPAIR-2 c.955 + round-3).", + ) + parser.add_argument( + "--calibration-size", type=int, default=CALIBRATION_SIZE, + help="Train-tail window for per-fold bias estimation (default: 60).", + ) args = parser.parse_args() coins = args.coins @@ -414,11 +962,11 @@ def main() -> None: coins = LOCAL_COINS if args.skip_remote else COINS print(f"M17 HAR-LJ-Asym: coins={coins}, horizons={args.horizons}, " - f"seeds={args.seeds}, skip_remote={args.skip_remote}") + f"seeds={args.seeds}, skip_remote={args.skip_remote}, " + f"debias={args.debias}, calibration_size={args.calibration_size}") t0 = time.time() - # Load hourly returns panel = _load_panel(skip_remote=args.skip_remote) available = [c for c in coins if c in panel] if not available: @@ -426,11 +974,9 @@ def main() -> None: return print(f"Panel loaded: {list(panel.keys())} ({len(panel[available[0]])} bars for {available[0]})") - # Compute daily components components = compute_daily_components(panel) print(f"Components computed for: {list(components.keys())}") - # Evaluate all combos rows: list[dict] = [] total = len(available) * len(args.horizons) * len(args.seeds) done = 0 @@ -439,7 +985,11 @@ def main() -> None: for seed in args.seeds: done += 1 print(f" [{done}/{total}] {coin} h={horizon} seed={seed}", end="", flush=True) - result = _eval_one_coin(coin, horizon, seed, components) + result = _eval_one_coin( + coin, horizon, seed, components, + debias=args.debias, + calibration_size=args.calibration_size, + ) if result is not None: rows.append(result) dm_h = result.get("dm_vs_har", {}).get("verdict", "?") @@ -451,7 +1001,6 @@ def main() -> None: elapsed = time.time() - t0 - # Aggregate agg = aggregate_verdicts(rows) output = { @@ -465,6 +1014,9 @@ def main() -> None: "fee_bps": FEE_BPS, "n_splits": N_SPLITS, "refit_every": REFIT_EVERY, + "debias_har": args.debias, + "calibration_size": args.calibration_size, + "calibration_protocol": "REPAIR-2 c.955 per-fold train-tail bias (no OOS target access)", }, "coins": available, "horizons": args.horizons, @@ -481,9 +1033,279 @@ def main() -> None: with open(out_path, "w", encoding="utf-8") as f: json.dump(output, f, indent=2, default=str) print(f"\nResults saved to {out_path}") + + # --- Concern #4: persist a manifest OUTSIDE the JSON results blob --- + # Manifest hash includes the entire run signature: rows + params + agg. + # Round-3 concern #6: cross-seed panel-hash consistency is meaningful + # WITHIN a (coin, horizon) group, not globally (different coins + # legitimately hash differently) -- surface the per-group bookkeeping. + ph_groups: dict[tuple, list[str]] = {} + for r in rows: + ph_groups.setdefault((r["coin"], r["horizon"]), []).append( + r.get("panel_hash", ""), + ) + panel_hash_per_coin_horizon = [ + { + "coin": c, + "horizon": h, + "panel_hash": hs[0] if hs else "", + "n_seed_rows": len(hs), + "consistent_across_seeds": len(set(hs)) <= 1, + } + for (c, h), hs in sorted(ph_groups.items()) + ] + # Round-4 concern (c): per-(coin, horizon) train/OOS bounds provenance, + # same surfacing pattern as panel_hash_per_coin_horizon -- the manifest + # says WHERE (index bounds) the forecasts/targets come from, alongside + # the existing value hashes. + bounds_groups: dict[tuple, list[dict]] = {} + for r in rows: + if r.get("bounds_train_test"): + bounds_groups.setdefault((r["coin"], r["horizon"]), []).append( + r["bounds_train_test"], + ) + bounds_per_coin_horizon = [ + { + "coin": c, + "horizon": h, + "bounds_train_test": bs[0], + "n_seed_rows": len(bs), + "consistent_across_seeds": ( + len({json.dumps(b, sort_keys=True) for b in bs}) <= 1 + ), + } + for (c, h), bs in sorted(bounds_groups.items()) + ] + manifest = { + "model": "M17_HAR_LJ_ASYM", + "params": output["params"], + "coins": available, + "horizons": args.horizons, + "seeds": args.seeds, + "elapsed_seconds": round(elapsed, 1), + "n_combos_evaluated": len(rows), + "bounds": { + "first_bar": ( + str(panel[available[0]].index[0].isoformat()) + if available and len(panel[available[0]]) else None + ), + "last_bar": ( + str(panel[available[0]].index[-1].isoformat()) + if available and len(panel[available[0]]) else None + ), + "n_bars_per_coin": { + c: int(len(panel[c])) for c in available if c in panel + }, + }, + "panel_hashes": sorted({r["panel_hash"] for r in rows if r.get("panel_hash")}), + "panel_hash_per_coin_horizon": panel_hash_per_coin_horizon, + "bounds_per_coin_horizon": bounds_per_coin_horizon, + "panel_hashes_consistent": ( + all(g["consistent_across_seeds"] for g in panel_hash_per_coin_horizon) + if panel_hash_per_coin_horizon else True + ), + "fc_hashes_per_coin_horizon": [ + { + "coin": r["coin"], "horizon": r["horizon"], "seed": r["seed"], + "fc_lj_hash": r["fc_lj_hash"], "fc_har_hash": r["fc_har_hash"], + "fc_m12_hash": r["fc_m12_hash"], "tgt_hash": r["tgt_hash"], + "err_lj_hash": r["err_lj_hash"], "err_har_hash": r["err_har_hash"], + "err_m12_hash": r["err_m12_hash"], "n_obs": r["n_obs"], + } + for r in rows + ], + "bit_identity_check": ( + "Bit-identity cross-seed only meaningful WITHIN the same (coin, " + "horizon, panel) tuple -- the seed does NOT change the OLS fit on " + "a fixed (X, y), so all seeds should produce identical forecasts " + "and DM verdicts (panel_hash + fc_*_hash consistent across seeds)." + ), + "edge_sigma_disposition": ( + "N/A. OLS on a deterministic (X, y) panel with fixed seeds is " + "bit-identical -- the panel_hash + per-row fc_hash/tgt_hash/" + "err_hash consistent across seeds {0, 7, 42, 99} is the " + "verifiable anchor. Multi-seed edge-σ is not applicable to " + "deterministic OLS; it applies to stochastic estimators (e.g. " + "neural nets with dropout, MCTS planners)." + ), + "concern_addressing": { + "concern_1_calibration_train_only": ( + "Per-fold bias estimated from train tail only via " + "_train_tail_bias() -- the bias estimate NEVER reads the OOS " + "target. This replaces the c.953 global tail-mean block that " + "leaked targets via mean(err[-60:]). Anti-leak test: " + "test_calibration_anti_leak_perturbation in test_har_lj_asym.py." + ), + "concern_2_har_not_double_calibrated": ( + "HAR receives calibrate_bias=True INSIDE walk_forward_har " + "(canonical train-tail bias). The post-walk-forward global " + "tail-mean block has been REMOVED -- mse_har_raw is now the " + "canonical HAR walk_forward output, not a post-corrected " + "double-calibrated value." + ), + "concern_3_dm_verdict_consistency": ( + "Each row carries the full DM verdict dict (dm_statistic, " + "p_value, mean_loss_diff, n_obs). The aggregated counts use " + "_coherent_beats() which REQUIRES (p_value < 0.05 AND " + "mean_loss_diff < 0) for BEATS, and (p_value < 0.05 AND " + "mean_loss_diff > 0) for BEATEN BY. The aggregator surfaces " + "p_values, dm_statistics, and mean_loss_diffs as lists per " + "(coin, horizon)." + ), + "concern_4_manifest_outside_git": ( + "scripts/results/manifest_m17_har_lj_asym.json is written " + "at every run with bounds (first/last bar per coin), panel " + "hashes, per-row forecast/target/error hashes, bit-identity " + "disposition, and edge-σ disposition." + ), + "concern_5_prev_valid": ( + "prev: MED/training #14561 (last MERGED training PR of this " + "lane, distinct from #14592)." + ), + "concern_1_round3_sign_and_per_fold": ( + "Bias convention is bias = mean(y_train_tail - yhat_train_tail); " + "corrected forecast = yhat + bias (the c.955 yhat - bias inverted " + "the sign and doubled the error). _eval_one_coin consumes the " + "per-fold-corrected forecasts_debiased array instead of a " + "post-walk-forward global-mean shift." + ), + "concern_2_round3_m12_calibrated": ( + "walk_forward_har_rv_j already exposes calibrate_bias (same " + "internal train-tail pattern as walk_forward_har); " + "_eval_one_coin now passes calibrate_bias=debias + " + "calibration_size so M12 is calibrated apples-to-apples." + ), + "concern_3_round3_har_raw_vs_debiased": ( + "mse_har_raw comes from an uncalibrated walk_forward_har pass; " + "mse_har_debiased from a separate calibrate_bias=True pass " + "(the DM leg). Two distinct legs -- the c.955 identity " + "mse_har_raw == mse_har_debiased is gone." + ), + "concern_4_round3_anti_leak_full_walk_forward": ( + "test_walk_forward_lj_asym_oos_target_invariance perturbs ALL " + "OOS targets through the full walk-forward and asserts " + "per_fold_bias and forecasts_debiased bit-identical." + ), + "concern_5_round3_live_run_deferred": ( + "c.953 live numbers are SUPERSEDED by the round-3 calibration " + "(see docs/M17_HAR_LJ_ASYM.md); the live BTC re-run is deferred " + "post-merge and gated by this manifest." + ), + "concern_6_round3_panel_hash_index": ( + "panel_hash = sha256(index_int64_ns || values_float64) on the " + "canonical 360-bar window; the manifest surfaces " + "panel_hash_per_coin_horizon (per-group hashes + consistency), " + "not a single collapsed global hash." + ), + "concern_a_round4_multifold_oos_invariance": ( + "test_walk_forward_lj_asym_oos_target_invariance_multi_fold " + "runs n_splits=3 with >=2 DISTINCT per-fold biases; shifting " + "only fold k's OOS target window leaves per_fold_bias[k] and " + "fold k's forecast slices bit-identical, with earlier folds " + "unchanged and later folds legitimately retraining on the " + "shifted rows (walk-forward expanding window, not leakage). " + "Per-fold train-tail shifts move only that fold's bias (>1.0)." + ), + "concern_c_round4_bounds_provenance": ( + "walk_forward_lj_asym surfaces bounds_train_test " + "{train_end_idx, oos_start_idx, oos_end_idx, n_train, n_oos} " + "(indices in X_all/merged-valid coordinates; train_end_idx = " + "n_folds * fold_size). _eval_one_coin, aggregate_verdicts " + "(bounds_train_test + bounds_consistent_across_seeds) and " + "this manifest (bounds_per_coin_horizon) relay it; " + "fc_lj_hash_per_fold is aligned with per_fold_bias. The " + "`if False else None` placeholder is removed." + ), + "concern_1_round5_per_fold_bounds_aligned": ( + "walk_forward_lj_asym now also surfaces per_fold_bounds -- a " + "list of {fold_idx, train_end_idx, oos_start_idx, oos_end_idx, " + "n_train, n_oos} dicts, one per fold, aligned 1-pour-1 with " + "per_fold_bias. _eval_one_coin relays per_fold_bounds per " + "(coin, horizon, seed); aggregate_verdicts surfaces " + "per_fold_bounds + per_fold_bounds_consistent_across_seeds per " + "(coin, horizon). The previous bounds_train_test only described " + "the last fold's cutoff; the forecasts/targets arrays are the " + "concatenation of 5 folds, so per-fold granularity is the " + "verifiable audit anchor." + ), + "concern_2_round5_content_hash_index_bornes_values": ( + "New helper _content_hash(index, bornes, values) digests " + "(int64 index bytes) || (int64 bornes bytes) || (float64 values " + "bytes) in that order. Surfaced as fc_content_hash_per_fold per " + "(coin, horizon, seed), parallel to fc_lj_hash_per_fold (values-" + "only). Mutation tests: test_content_hash_mutates_on_index_shift, " + "test_content_hash_mutates_on_bounds_change, " + "test_content_hash_mutates_on_value_change (test_har_lj_asym.py)." + ), + "concern_3_round5_durable_anchor_no_private_data": ( + "main() writes scripts/results/run_anchor_m17_har_lj_asym.json " + "with the SHA-256 + relative locus of both the manifest and " + "the results JSON (the files themselves are gitignored). The " + "anchor is the durable reference; the underlying data stays " + "local. Reviewer verification = sha256sum on the loci listed " + "in run_anchor.<...>.sha256 fields -- no private data leaves " + "the worktree." + ), + }, + "manifest_sha256": hashlib.sha256( + json.dumps(output, sort_keys=True, default=str).encode("utf-8") + ).hexdigest(), + } + manifest_path = RESULTS_DIR / "manifest_m17_har_lj_asym.json" + with open(manifest_path, "w", encoding="utf-8") as f: + json.dump(manifest, f, indent=2, default=str) + print(f"Manifest saved to {manifest_path}") + + # Round-5 concern (3): durable anchor for the run WITHOUT publishing + # the underlying private data. The manifest + results JSON files are + # gitignored (.gitignore:51:results/), so the data stays local; the + # SHA-256 of each artefact + the relative locus (path from the repo + # root) is the durable anchor that a reviewer can verify by re-running + # ``sha256sum`` on the same files post-merge (no data shipped). + def _sha256(p: "Path") -> str: + return hashlib.sha256(p.read_bytes()).hexdigest() + run_anchor = { + "model": "M17_HAR_LJ_ASYM", + "run_signature": { + "n_combos_evaluated": len(rows), + "elapsed_seconds": round(elapsed, 1), + "coins": available, + "horizons": args.horizons, + "seeds": args.seeds, + "skip_remote": args.skip_remote, + "debias": args.debias, + "calibration_size": args.calibration_size, + }, + "manifest_anchor": { + "locus": str(manifest_path), + "sha256": _sha256(manifest_path), + "bytes": int(manifest_path.stat().st_size), + }, + "results_anchor": { + "locus": str(out_path), + "sha256": _sha256(out_path), + "bytes": int(out_path.stat().st_size), + }, + "anchor_protocol": ( + "Both artefacts are gitignored (results/ in the ML-Training-Pipeline " + "subtree); the SHA-256 + relative locus is the durable reference. " + "Reviewer verification: `sha256sum ` against `anchor.<...>.sha256`." + ), + } + anchor_path = RESULTS_DIR / "run_anchor_m17_har_lj_asym.json" + with open(anchor_path, "w", encoding="utf-8") as f: + json.dump(run_anchor, f, indent=2, default=str) + print(f"Run anchor saved to {anchor_path}") + # Stdout summary: short keys only, no path content -- the user / + # reviewer can copy these SHA-256 prefixes to compare against the + # files on disk without polluting logs with full paths. + print(f" manifest_sha256[:16] = {run_anchor['manifest_anchor']['sha256'][:16]}") + print(f" results_sha256[:16] = {run_anchor['results_anchor']['sha256'][:16]}") + print(f" manifest_locus = {run_anchor['manifest_anchor']['locus']}") + print(f" results_locus = {run_anchor['results_anchor']['locus']}") + print(f"Total: {len(rows)} combos evaluated in {elapsed:.1f}s") - # Summary if agg: print("\n=== Aggregated Results ===") for a in agg: diff --git a/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/scripts/tests/test_har_lj_asym.py b/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/scripts/tests/test_har_lj_asym.py new file mode 100644 index 0000000000..4156ef7cdb --- /dev/null +++ b/MyIA.AI.Notebooks/QuantConnect/ML-Training-Pipeline/scripts/tests/test_har_lj_asym.py @@ -0,0 +1,1348 @@ +"""Tests for the M17 HAR-LJ-Asym walk-forward evaluation. + +REPAIR-2 (c.955): the previous version of these tests assumed the c.953 +"symmetric" calibration (post-walk-forward global tail-mean block) — that +calibration protocol read OOS targets via ``mean(err[-60:])``, which is the +fuite that the preflight po-2025 re-review (head 4cc2262b) flagged. The +tests now validate the per-fold train-tail bias protocol that REPAIR-2 +introduces, plus the new coherence requirements (DM verdict <-> p_value + +mean_loss_diff), the manifest hashes, and the bit-identity anchor across +seeds. + +Concerns addressed (verbatim from the c.955 re-review): + #1 — calibration must NOT read OOS targets (anti-leak test below). + #2 — HAR must NOT be double-calibrated. + #3 — DM verdict components must be coherent (BEATS => p<0.05 AND diff<0). + #4 — manifest hashes (forecasts/targets/errors) must be emitted. + +ROUND-3 (preflight po-2025 adjoint re-review, head b974f2721, DM +msg-20260904T141944): + #1 — sign of the bias correction (+, not -) + per-fold constant shift. + #2 — walk_forward_har_rv_j receives calibrate_bias=debias (M12 calibrated). + #3 — mse_har_raw (uncalibrated leg) != mse_har_debiased (calibrated leg). + #4 — full-walk-forward OOS-target invariance test. + #6 — panel_hash covers the index in addition to the values. + +ROUND-4 (adjoint re-review, DM msg-20260905T001520, 3/6 PARTIAL): + (a) multi-fold OOS-target invariance — n_splits=3, DISTINCT per-fold + biases, per-fold assertions (the round-3 single-fold test stays as a + smoke test). + (c) bounds provenance — bounds_train_test surfaced by + walk_forward_lj_asym / _eval_one_coin / aggregate_verdicts, with + fc_lj_hash_per_fold aligned with per_fold_bias. +""" + +from __future__ import annotations + +import hashlib +import json +from pathlib import Path +import sys + +import numpy as np +import pandas as pd +import pytest + +SCRIPTS_DIR = Path(__file__).resolve().parent.parent +sys.path.insert(0, str(SCRIPTS_DIR)) + +from har_lj_asym import ( # noqa: E402 (sys.path mutation before import) + _train_tail_bias, + aggregate_verdicts, + HARLJAsymModel, +) + + +# --------------------------------------------------------------------------- +# Fixtures +# --------------------------------------------------------------------------- + + +@pytest.fixture +def synthetic_lj_components() -> dict[str, dict[str, pd.Series]]: + """Synthetic RV + jump + semivariance components for 360 days.""" + rng = np.random.default_rng(23) + index = pd.date_range("2020-01-01", periods=360, freq="D") + + log_rv = np.empty(len(index)) + log_rv[0] = -7.5 + for i in range(1, len(index)): + log_rv[i] = -1.2 + 0.84 * log_rv[i - 1] + rng.normal(0.0, 0.12) + rv = pd.Series(np.exp(log_rv), index=index, name="rv") + + downside_share = np.clip( + 0.55 + rng.normal(0.0, 0.05, len(index)), 0.2, 0.8, + ) + rv_neg = pd.Series( + rv.to_numpy() * downside_share, index=index, name="rv_neg", + ) + rv_pos = pd.Series( + rv.to_numpy() * (1.0 - downside_share), index=index, name="rv_pos", + ) + jumps = pd.Series( + np.clip(rv.to_numpy() * rng.uniform(0.0, 0.15, len(index)), 0, None), + index=index, name="rv_j", + ) + rv_c = pd.Series( + np.clip(rv.to_numpy() - jumps.to_numpy(), 1e-12, None), + index=index, name="rv_c", + ) + return { + "BTC-USD": { + "rv": rv, + "rv_neg": rv_neg, + "rv_pos": rv_pos, + "rv_c": rv_c, + "rv_j": jumps, + }, + } + + +# --------------------------------------------------------------------------- +# Concern #1 — anti-leak test for the per-fold train-tail bias estimator +# --------------------------------------------------------------------------- + + +def test_calibration_anti_leak_perturbation(): + """Concern #1 (c.955): perturbing the OOS targets must NOT change the + per-fold bias estimate. The bias is computed from the train tail ONLY; + OOS target perturbation is a no-op for ``_train_tail_bias``. + + This is the apples-to-apples test demanded by #14584 disposition #1: + perturb the targets (y_all), refit the model, and verify the bias + estimate from train tail is unchanged across the perturbation. + """ + rng = np.random.default_rng(42) + n_train, n_features = 200, 6 + X_train = rng.normal(0.0, 1.0, (n_train, n_features)) + y_train = -1.5 + 0.5 * X_train[:, 0] + 0.3 * X_train[:, 1] + rng.normal(0.0, 0.1, n_train) + + model = HARLJAsymModel().fit(X_train, y_train) + + bias_unperturbed = _train_tail_bias(model, X_train, y_train, calibration_size=60) + # Perturb the LAST 60 train targets (the calibration tail itself -- still + # train data, no OOS access). The bias estimate must change because we + # are perturbing the tail we sample from. This is the **expected** behavior. + y_train_perturbed_tail = y_train.copy() + y_train_perturbed_tail[-60:] += 5.0 # add a large constant shift + bias_perturbed_tail = _train_tail_bias(model, X_train, y_train_perturbed_tail, calibration_size=60) + assert abs(bias_perturbed_tail - bias_unperturbed) > 1.0, ( + "Perturbing the train tail SHOULD shift the bias estimate -- this is " + "the expected behavior. The bias estimator reads from the train tail " + "and is sensitive to train perturbations by design." + ) + + # Now perturb the FIRST 140 train targets (NOT in the calibration tail). + # The bias estimate from the LAST 60 must be unchanged. + y_train_perturbed_head = y_train.copy() + y_train_perturbed_head[:140] += 5.0 + bias_perturbed_head = _train_tail_bias(model, X_train, y_train_perturbed_head, calibration_size=60) + np.testing.assert_allclose( + bias_perturbed_head, bias_unperturbed, rtol=1e-12, + err_msg=( + "Perturbing the FIRST 140 train targets (outside the calibration " + "tail) MUST NOT shift the bias estimate from the last 60." + ), + ) + + +# --------------------------------------------------------------------------- +# Concern #2 — HAR is not double-calibrated (calibrate_bias propagated once) +# --------------------------------------------------------------------------- + + +def test_walk_forward_har_called_with_calibrate_bias_when_debias( + synthetic_lj_components, monkeypatch, +): + """Concern #2 (c.955) + round-3 concern #3: when debias=True, + walk_forward_har is called TWICE -- first uncalibrated (mse_har_raw leg), + then with calibrate_bias=True (DM + mse_har_debiased leg). No post-walk- + forward second correction is applied on top of the calibrated leg.""" + from har_lj_asym import _eval_one_coin + + captured: list = [] + + def spy_walk_forward_har(rv, horizon, *args, **kwargs): + captured.append( + (kwargs.get("calibrate_bias"), kwargs.get("calibration_size")), + ) + n = 100 + idx = rv.index[-n:] + return { + "forecasts": pd.Series(np.zeros(n), index=idx, name="fc"), + "aggregate_mse_logrv": 0.0, + } + + monkeypatch.setattr("har_lj_asym.walk_forward_har", spy_walk_forward_har) + + _eval_one_coin( + "BTC-USD", horizon=1, seed=0, components=synthetic_lj_components, + debias=True, calibration_size=60, + ) + + # Raw leg first (calibrate_bias falsy), calibrated DM leg second. + assert [c[0] for c in captured] == [False, True] + assert captured[1][1] == 60 + + +def test_walk_forward_har_called_with_calibrate_bias_false_when_no_debias( + synthetic_lj_components, monkeypatch, +): + """When debias=False, walk_forward_har must receive calibrate_bias=False + exactly once (no calibrated second leg is needed).""" + from har_lj_asym import _eval_one_coin + + captured: list = [] + + def spy_walk_forward_har(rv, horizon, *args, **kwargs): + captured.append(kwargs.get("calibrate_bias")) + n = 100 + idx = rv.index[-n:] + return { + "forecasts": pd.Series(np.zeros(n), index=idx, name="fc"), + "aggregate_mse_logrv": 0.0, + } + + monkeypatch.setattr("har_lj_asym.walk_forward_har", spy_walk_forward_har) + + _eval_one_coin( + "BTC-USD", horizon=1, seed=0, components=synthetic_lj_components, + ) + + assert captured == [False] + + +# --------------------------------------------------------------------------- +# Concern #3 — DM verdict coherence: BEATS => p<0.05 AND mean_loss_diff<0 +# --------------------------------------------------------------------------- + + +def test_aggregate_coherent_beats_requires_p_lt_alpha_and_diff_lt_zero(): + """Concern #3: a BEATS verdict must imply p_value < 0.05 AND + mean_loss_diff < 0. If a row reports BEATS but p_value >= 0.05, it is + mis-classified upstream and must NOT be counted as a win. + """ + rows = [ + { + "coin": "BTC-USD", "horizon": 1, "seed": s, + "mse_logrv": 0.84, "mse_har_raw": 1.0, "mse_har_debiased": 1.0, + "mse_m12": 1.1, "bias_lj": 0.0, "bias_har": 0.0, "bias_m12": 0.0, + "var_lj": 0.84, "var_har": 1.0, "var_m12": 1.1, + "sharpe": np.nan, "kelly_active_pct": 0.5, + # Mis-classified: BEATS verdict but p_value >= 0.05. + "dm_vs_har": { + "verdict": "BEATS baseline", "p_value": 0.83, + "mean_loss_diff": -0.01, "dm_statistic": -2.0, "n_obs": 100, + "lag": 4, "hac_variance": 0.001, "significant_at": 0.05, + }, + "dm_vs_m12": { + "verdict": "INCONCLUSIVE", "p_value": 0.5, + "mean_loss_diff": -0.005, "dm_statistic": -1.0, "n_obs": 100, + "lag": 4, "hac_variance": 0.001, "significant_at": 0.05, + }, + "panel_hash": "deadbeef", + "fc_lj_hash": "a", "fc_har_hash": "b", "fc_m12_hash": "c", + "tgt_hash": "d", "err_lj_hash": "e", "err_har_hash": "f", + "err_m12_hash": "g", "n_obs": 100, "edge_sigma_applicable": False, + } + for s in (0, 7, 42, 99) + ] + agg = aggregate_verdicts(rows)[0] + # The BEATS verdict with p_value=0.83 must NOT be counted as a win. + assert agg["dm_vs_har_wins"] == 0, ( + f"Coherence violation: BEATS verdict with p_value >= 0.05 should NOT " + f"count as a win. Got {agg['dm_vs_har_wins']} wins." + ) + + +def test_aggregate_coherent_beats_counts_when_p_lt_alpha_and_diff_lt_zero(): + """Conversely, when the coherence holds (p<0.05 AND diff<0), BEATS + must be counted as a win.""" + rows = [ + { + "coin": "BTC-USD", "horizon": 1, "seed": s, + "mse_logrv": 0.84, "mse_har_raw": 1.0, "mse_har_debiased": 1.0, + "mse_m12": 1.1, "bias_lj": 0.0, "bias_har": 0.0, "bias_m12": 0.0, + "var_lj": 0.84, "var_har": 1.0, "var_m12": 1.1, + "sharpe": np.nan, "kelly_active_pct": 0.5, + # Coherent: BEATS verdict + p_value < 0.05 + diff < 0. + "dm_vs_har": { + "verdict": "BEATS baseline", "p_value": 0.01, + "mean_loss_diff": -0.05, "dm_statistic": -3.0, "n_obs": 100, + "lag": 4, "hac_variance": 0.001, "significant_at": 0.05, + }, + "dm_vs_m12": { + "verdict": "INCONCLUSIVE", "p_value": 0.5, + "mean_loss_diff": -0.005, "dm_statistic": -1.0, "n_obs": 100, + "lag": 4, "hac_variance": 0.001, "significant_at": 0.05, + }, + "panel_hash": "deadbeef", + "fc_lj_hash": "a", "fc_har_hash": "b", "fc_m12_hash": "c", + "tgt_hash": "d", "err_lj_hash": "e", "err_har_hash": "f", + "err_m12_hash": "g", "n_obs": 100, "edge_sigma_applicable": False, + } + for s in (0, 7, 42, 99) + ] + agg = aggregate_verdicts(rows)[0] + assert agg["dm_vs_har_wins"] == 4 + + +def test_aggregate_surfaces_dm_components_per_horizon(): + """Concern #3: aggregated DM components (p_values list, mean_loss_diffs + list, dm_statistics list) must be surfaced for auditability.""" + rows = [ + { + "coin": "BTC-USD", "horizon": 1, "seed": s, + "mse_logrv": 0.84, "mse_har_raw": 1.0, "mse_har_debiased": 1.0, + "mse_m12": 1.1, "bias_lj": 0.0, "bias_har": 0.0, "bias_m12": 0.0, + "var_lj": 0.84, "var_har": 1.0, "var_m12": 1.1, + "sharpe": np.nan, "kelly_active_pct": 0.5, + "dm_vs_har": { + "verdict": "INCONCLUSIVE", "p_value": 0.01 * (1 + s), + "mean_loss_diff": -0.01, "dm_statistic": -1.0 * s, "n_obs": 100, + "lag": 4, "hac_variance": 0.001, "significant_at": 0.05, + }, + "dm_vs_m12": { + "verdict": "INCONCLUSIVE", "p_value": 0.5, + "mean_loss_diff": -0.005, "dm_statistic": -1.0, "n_obs": 100, + "lag": 4, "hac_variance": 0.001, "significant_at": 0.05, + }, + "panel_hash": "deadbeef", + "fc_lj_hash": "a", "fc_har_hash": "b", "fc_m12_hash": "c", + "tgt_hash": "d", "err_lj_hash": "e", "err_har_hash": "f", + "err_m12_hash": "g", "n_obs": 100, "edge_sigma_applicable": False, + } + for s in (0, 7, 42, 99) + ] + agg = aggregate_verdicts(rows)[0] + har_components = agg["dm_vs_har_components"] + # Test data used p_value = 0.01 * (1 + s) with s in (0, 7, 42, 99) and + # dm_statistic = -1.0 * s. Assert on the per-row invariants + # (mean_loss_diff constant, dm_statistic follows -1.0 * s) and on the + # median (numpy median of even N is the mean of the two center values). + assert har_components["mean_loss_diffs"] == [-0.01] * 4 + assert har_components["dm_statistics"] == [0.0, -7.0, -42.0, -99.0] + sorted_p = sorted([0.01, 0.01 * 8, 0.01 * 43, 1.0]) + expected_median = (sorted_p[1] + sorted_p[2]) / 2 + assert har_components["p_value_median"] == pytest.approx(expected_median) + + +# --------------------------------------------------------------------------- +# Concern #4 — manifest hashes (forecasts/targets/errors) emitted per row +# --------------------------------------------------------------------------- + + +def test_eval_one_coin_emits_manifest_hashes( + synthetic_lj_components, monkeypatch, +): + """Concern #4: each row must carry forecast/target/error hashes for the + manifest audit (panel_hash + fc_*_hash + tgt_hash + err_*_hash).""" + from har_lj_asym import _eval_one_coin + + def fake_walk_forward_har(rv, horizon, *args, **kwargs): + n = 100 + idx = rv.index[-n:] + return { + "forecasts": pd.Series( + np.log(rv.iloc[-n:].to_numpy()) - 0.05, + index=idx, name="fc_har", + ), + "aggregate_mse_logrv": 0.9, + } + + monkeypatch.setattr("har_lj_asym.walk_forward_har", fake_walk_forward_har) + + row = _eval_one_coin( + "BTC-USD", horizon=1, seed=0, components=synthetic_lj_components, + debias=True, calibration_size=60, + ) + assert row is not None + # All manifest fields must be present and look like 16-hex-char SHA prefixes. + for field in ( + "panel_hash", "fc_lj_hash", "fc_har_hash", "fc_m12_hash", + "tgt_hash", "err_lj_hash", "err_har_hash", "err_m12_hash", + ): + v = row[field] + assert isinstance(v, str) + assert len(v) == 16 + int(v, 16) # hex parseable + + # Edge-sigma disposition is N/A (deterministic OLS). + assert row["edge_sigma_applicable"] is False + assert row["n_obs"] >= 10 + + +# --------------------------------------------------------------------------- +# Bit-identity anchor (c.953 sustained): panel_hash consistent across seeds +# --------------------------------------------------------------------------- + + +def test_panel_hash_consistent_across_seeds( + synthetic_lj_components, monkeypatch, +): + """Concern #3 (c.953): panel_hash on the canonical 360-bar RV window is + identical across seeds {0, 7, 42, 99}. Deterministic OLS guarantee.""" + from har_lj_asym import _eval_one_coin + + def fake_walk_forward_har(rv, horizon, *args, **kwargs): + n = 200 + idx = rv.index[-n:] + return { + "forecasts": pd.Series( + np.log(rv.iloc[-n:].to_numpy()) - 0.05, + index=idx, name="fc_har", + ), + "aggregate_mse_logrv": 0.9, + } + + monkeypatch.setattr("har_lj_asym.walk_forward_har", fake_walk_forward_har) + + rows = [] + for seed in (0, 7, 42, 99): + r = _eval_one_coin( + "BTC-USD", horizon=1, seed=seed, + components=synthetic_lj_components, + debias=True, calibration_size=60, + ) + assert r is not None + rows.append(r) + + panel_hashes = {r["panel_hash"] for r in rows} + assert len(panel_hashes) == 1, f"panel hashes diverge across seeds: {panel_hashes}" + + agg = aggregate_verdicts(rows)[0] + assert agg["panel_hashes_consistent"] is True + assert agg["panel_hash"] != "" + + +# --------------------------------------------------------------------------- +# Per-fold bias estimates surfaced for audit (concern #1) +# --------------------------------------------------------------------------- + + +def test_walk_forward_lj_asym_returns_per_fold_bias(): + """Concern #1 acceptance: walk_forward_lj_asym must surface + ``per_fold_bias`` (list of train-tail bias estimates, one per fold). + Round-3 note: ``_eval_one_coin`` consumes these PER FOLD through + ``forecasts_debiased`` (no global-mean aggregation anymore). + """ + from har_lj_asym import walk_forward_lj_asym + + rng = np.random.default_rng(0) + index = pd.date_range("2020-01-01", periods=400, freq="D") + log_rv = np.cumsum(rng.normal(0.0, 0.1, 400)) - 5.0 + rv = pd.Series(np.exp(log_rv), index=index, name="rv") + rv_neg = pd.Series(rv.to_numpy() * 0.5, index=index) + rv_pos = pd.Series(rv.to_numpy() * 0.5, index=index) + rv_c = pd.Series(rv.to_numpy() * 0.8, index=index) + rv_j = pd.Series(rv.to_numpy() * 0.2, index=index) + components = {"BTC-USD": { + "rv": rv, "rv_neg": rv_neg, "rv_pos": rv_pos, "rv_c": rv_c, "rv_j": rv_j, + }} + + out = walk_forward_lj_asym( + rv, rv_neg, rv_pos, rv_c, rv_j, horizon=1, seed=0, + debias=True, calibration_size=60, + ) + assert out["forecasts"] + assert len(out["per_fold_bias"]) > 0 + # When debias=True, forecasts_debiased must be present and len == forecasts. + assert out["forecasts_debiased"] + assert len(out["forecasts_debiased"]) == len(out["forecasts"]) + # Per-fold bias is a single scalar per fold; all entries should be finite floats. + for b in out["per_fold_bias"]: + assert np.isfinite(b) + + +# --------------------------------------------------------------------------- +# Manifest path + shape (concern #4 acceptance) +# --------------------------------------------------------------------------- + + +def test_manifest_path_constants(): + """Concern #4: the manifest path is `scripts/results/manifest_m17_har_lj_asym.json` + next to the main results JSON.""" + from har_lj_asym import RESULTS_DIR + assert (RESULTS_DIR / "manifest_m17_har_lj_asym.json").parent == RESULTS_DIR + + +# --------------------------------------------------------------------------- +# Backward compat: the c.953 invariant ``mse = bias^2 + var`` still holds +# --------------------------------------------------------------------------- + + +def test_mse_decomposition_equals_empirical_mean_squared_error( + synthetic_lj_components, monkeypatch, +): + """The c.953 invariant ``bias^2 + var(ddof=0) == mean(err**2)`` continues + to hold under the new protocol (concern #1 acceptance, sustained).""" + from har_lj_asym import _eval_one_coin + + def fake_walk_forward_har(rv, horizon, *args, **kwargs): + n = 200 + idx = rv.index[-n:] + return { + "forecasts": pd.Series(np.zeros(n), index=idx, name="fc_har"), + "aggregate_mse_logrv": 0.0, + } + + monkeypatch.setattr("har_lj_asym.walk_forward_har", fake_walk_forward_har) + + row = _eval_one_coin( + "BTC-USD", horizon=1, seed=0, components=synthetic_lj_components, + debias=False, calibration_size=60, + ) + assert row is not None + + np.testing.assert_allclose( + row["bias_lj"] ** 2 + row["var_lj"], row["mse_logrv"], rtol=1e-9, + ) + np.testing.assert_allclose( + row["bias_har"] ** 2 + row["var_har"], row["mse_har_raw"], rtol=1e-9, + ) + np.testing.assert_allclose( + row["bias_m12"] ** 2 + row["var_m12"], row["mse_m12"], rtol=1e-9, + ) + + # Probe: err=[0,1] -> bias=0.5, var(ddof=0)=0.25, MSE=0.5 + err_probe = np.array([0.0, 1.0]) + assert np.var(err_probe, ddof=0) == pytest.approx(0.25) + assert np.mean(err_probe ** 2) == pytest.approx(0.5) + + +def test_mse_har_debiased_is_nan_when_debias_false( + synthetic_lj_components, monkeypatch, +): + """c.953 invariant sustained: mse_har_debiased is NaN when debias=False.""" + from har_lj_asym import _eval_one_coin + + def fake_walk_forward_har(rv, horizon, *args, **kwargs): + n = 100 + idx = rv.index[-n:] + return { + "forecasts": pd.Series(np.zeros(n), index=idx, name="fc"), + "aggregate_mse_logrv": 0.0, + } + + monkeypatch.setattr("har_lj_asym.walk_forward_har", fake_walk_forward_har) + + row = _eval_one_coin( + "BTC-USD", horizon=1, seed=0, components=synthetic_lj_components, + debias=False, + ) + assert row is not None + assert np.isnan(row["mse_har_debiased"]) + assert not np.isnan(row["mse_har_raw"]) + assert row["mse_har_raw"] >= 0.0 + + +# --------------------------------------------------------------------------- +# Aggregation: var_ratio + DM counts (sustained from c.953) +# --------------------------------------------------------------------------- + + +def _row( + coin: str, horizon: int, seed: int, + mse_logrv: float, mse_har: float, mse_m12: float, + bias_lj: float, bias_har: float, bias_m12: float, + var_lj: float, var_har: float, var_m12: float, + dm_har: str, dm_m12: str, + sharpe: float = np.nan, + mse_har_debiased: float | None = None, + panel_hash: str = "deadbeef", + p_value_har: float = 0.5, p_value_m12: float = 0.5, +) -> dict: + return { + "coin": coin, "horizon": horizon, "seed": seed, + "mse_logrv": mse_logrv, + "mse_har_raw": mse_har, + "mse_har_debiased": (mse_har_debiased if mse_har_debiased is not None else mse_har), + "mse_m12": mse_m12, + "bias_lj": bias_lj, "bias_har": bias_har, "bias_m12": bias_m12, + "var_lj": var_lj, "var_har": var_har, "var_m12": var_m12, + "sharpe": sharpe, "kelly_active_pct": 0.5, + "dm_vs_har": { + "verdict": dm_har, "p_value": p_value_har, + "mean_loss_diff": -0.01 if dm_har == "BEATS baseline" else 0.0, + "dm_statistic": -2.0, "n_obs": 100, "lag": 4, + "hac_variance": 0.001, "significant_at": 0.05, + }, + "dm_vs_m12": { + "verdict": dm_m12, "p_value": p_value_m12, + "mean_loss_diff": -0.01 if dm_m12 == "BEATS baseline" else 0.0, + "dm_statistic": -2.0, "n_obs": 100, "lag": 4, + "hac_variance": 0.001, "significant_at": 0.05, + }, + "panel_hash": panel_hash, + "fc_lj_hash": "a", "fc_har_hash": "b", "fc_m12_hash": "c", + "tgt_hash": "d", "err_lj_hash": "e", "err_har_hash": "f", + "err_m12_hash": "g", "n_obs": 100, "edge_sigma_applicable": False, + } + + +def test_aggregate_var_ratio_lj_over_har(): + """Sustained from c.953: var_ratio aggregates per-seed var_lj / mean of var_har.""" + rows = [ + _row( + "BTC-USD", 1, seed, + mse_logrv=0.84, mse_har=1.08, mse_m12=1.13, + bias_lj=0.025, bias_har=-0.002, bias_m12=-0.244, + var_lj=0.839, var_har=1.078, var_m12=1.072, + dm_har="BEATS baseline", dm_m12="BEATS baseline", + p_value_har=0.01, p_value_m12=0.01, + ) + for seed in (0, 7, 42, 99) + ] + agg = aggregate_verdicts(rows)[0] + assert agg["avg_var_lj"] == pytest.approx(0.839) + assert agg["avg_var_har"] == pytest.approx(1.078) + assert agg["var_ratio_lj_over_har"] == pytest.approx(0.839 / 1.078, rel=1e-3) + assert agg["dm_vs_har_wins"] == 4 + assert agg["dm_vs_m12_wins"] == 4 + + +def test_aggregate_var_ratio_handles_zero_baseline_safely(): + """Sustained from c.953: NaN-safe division when var_har is exactly 0.""" + rows = [ + _row( + "BTC-USD", 1, seed, + mse_logrv=0.0, mse_har=0.0, mse_m12=0.0, + bias_lj=0.0, bias_har=0.0, bias_m12=0.0, + var_lj=1.0, var_har=0.0, var_m12=0.0, + dm_har="INCONCLUSIVE", dm_m12="INCONCLUSIVE", + ) + for seed in (0, 7, 42, 99) + ] + agg = aggregate_verdicts(rows)[0] + assert np.isnan(agg["var_ratio_lj_over_har"]) + + +def test_aggregate_dm_verdict_counts_separated_per_baseline(): + """Sustained from c.953: counts tracked independently for HAR and M12.""" + rows = [] + for seed, dm_h in zip( + (0, 7, 42, 99), + ("BEATS baseline", "BEATS baseline", "INCONCLUSIVE", "INCONCLUSIVE"), + ): + rows.append(_row( + "BTC-USD", 5, seed, + mse_logrv=0.40, mse_har=0.38, mse_m12=0.52, + bias_lj=0.0, bias_har=0.0, bias_m12=-0.38, + var_lj=0.40, var_har=0.38, var_m12=0.37, + dm_har=dm_h, dm_m12="BEATS baseline", + p_value_har=0.01, p_value_m12=0.01, + )) + agg = aggregate_verdicts(rows)[0] + assert agg["dm_vs_har_wins"] == 2 + assert agg["dm_vs_har_total"] == 4 + assert agg["dm_vs_m12_wins"] == 4 + assert agg["dm_vs_m12_total"] == 4 + + +def test_aggregate_seeds_preserved_for_audit(): + """Sustained from c.953: the aggregator surfaces the seeds list verbatim.""" + rows = [ + _row( + "BTC-USD", 1, seed, + mse_logrv=0.8, mse_har=1.0, mse_m12=1.1, + bias_lj=0.0, bias_har=0.0, bias_m12=0.0, + var_lj=0.8, var_har=1.0, var_m12=1.1, + dm_har="INCONCLUSIVE", dm_m12="INCONCLUSIVE", + ) + for seed in (0, 7, 42, 99) + ] + agg = aggregate_verdicts(rows)[0] + assert agg["seeds"] == [0, 7, 42, 99] + assert agg["n_seeds"] == 4 + + +# --------------------------------------------------------------------------- +# ROUND-3 concern #1 — sign of the correction + per-fold constant shift +# --------------------------------------------------------------------------- + + +def test_forecasts_debiased_is_forecasts_plus_per_fold_bias(): + """Round-3 concern #1: ``forecasts_debiased`` must equal + ``forecasts + per_fold_bias[f]`` on each fold slice (the bias is ADDED, + per the sign convention ``bias = mean(y_tail - yhat_tail)``), and the + bias must be non-trivial on this synthetic panel so the test cannot pass + vacuously at bias == 0. + """ + from har_lj_asym import walk_forward_lj_asym + + rng = np.random.default_rng(0) + index = pd.date_range("2020-01-01", periods=400, freq="D") + log_rv = np.cumsum(rng.normal(0.0, 0.1, 400)) - 5.0 + rv = pd.Series(np.exp(log_rv), index=index, name="rv") + rv_neg = pd.Series(rv.to_numpy() * 0.5, index=index) + rv_pos = pd.Series(rv.to_numpy() * 0.5, index=index) + rv_c = pd.Series(rv.to_numpy() * 0.8, index=index) + rv_j = pd.Series(rv.to_numpy() * 0.2, index=index) + + out = walk_forward_lj_asym( + rv, rv_neg, rv_pos, rv_c, rv_j, horizon=1, seed=0, + debias=True, calibration_size=60, + ) + assert out["forecasts"] + biases = out["per_fold_bias"] + assert len(biases) > 0 + # Non-trivial bias on this synthetic panel (guards against a vacuous + # test where +bias and -bias are indistinguishable at bias == 0). + assert float(np.max(np.abs(biases))) > 1e-6 + + fc = np.asarray(out["forecasts"], dtype=float) + fcd = np.asarray(out["forecasts_debiased"], dtype=float) + shift = fcd - fc + fold_size = len(fc) // len(biases) + assert fold_size * len(biases) == len(fc) + for k, b in enumerate(biases): + np.testing.assert_allclose( + shift[k * fold_size:(k + 1) * fold_size], b, + rtol=1e-12, atol=1e-12, + err_msg=f"fold {k}: debiased forecasts must equal forecasts + bias", + ) + + +# --------------------------------------------------------------------------- +# ROUND-3 concern #2 — M12 (walk_forward_har_rv_j) is calibrated when debias +# --------------------------------------------------------------------------- + + +def test_walk_forward_har_rv_j_receives_calibrate_bias_when_debias( + synthetic_lj_components, monkeypatch, +): + """Round-3 concern #2: when debias=True, walk_forward_har_rv_j must be + called with calibrate_bias=True + the same calibration_size (apples-to- + apples M12 baseline, internally calibrated like HAR and LJ).""" + import m12_har_rv_j + from har_lj_asym import _eval_one_coin + + captured: dict = {} + + def spy_walk_forward_har_rv_j(rv, jumps, horizon, *args, **kwargs): + captured["calibrate_bias"] = kwargs.get("calibrate_bias") + captured["calibration_size"] = kwargs.get("calibration_size") + n = 100 + idx = rv.index[-n:] + return { + "forecasts": pd.Series(np.zeros(n), index=idx, name="fc"), + "aggregate_mse_logrv": 0.0, + } + + monkeypatch.setattr( + m12_har_rv_j, "walk_forward_har_rv_j", spy_walk_forward_har_rv_j, + ) + + row = _eval_one_coin( + "BTC-USD", horizon=1, seed=0, components=synthetic_lj_components, + debias=True, calibration_size=60, + ) + assert row is not None + assert captured["calibrate_bias"] is True + assert captured["calibration_size"] == 60 + + +def test_walk_forward_har_rv_j_receives_calibrate_bias_false_when_no_debias( + synthetic_lj_components, monkeypatch, +): + """Round-3 concern #2: when debias=False, walk_forward_har_rv_j must + receive calibrate_bias=False (no calibration anywhere).""" + import m12_har_rv_j + from har_lj_asym import _eval_one_coin + + captured: dict = {} + + def spy_walk_forward_har_rv_j(rv, jumps, horizon, *args, **kwargs): + captured["calibrate_bias"] = kwargs.get("calibrate_bias") + n = 100 + idx = rv.index[-n:] + return { + "forecasts": pd.Series(np.zeros(n), index=idx, name="fc"), + "aggregate_mse_logrv": 0.0, + } + + monkeypatch.setattr( + m12_har_rv_j, "walk_forward_har_rv_j", spy_walk_forward_har_rv_j, + ) + + row = _eval_one_coin( + "BTC-USD", horizon=1, seed=0, components=synthetic_lj_components, + ) + assert row is not None + assert captured["calibrate_bias"] is False + + +# --------------------------------------------------------------------------- +# ROUND-3 concern #3 — mse_har_raw != mse_har_debiased with calibrate-aware +# fixtures (the two HAR legs must be genuinely distinct forecasts) +# --------------------------------------------------------------------------- + + +def test_mse_har_raw_and_debiased_distinct_when_debias( + synthetic_lj_components, monkeypatch, +): + """Round-3 concern #3: with a HAR fixture whose forecasts DEPEND on + calibrate_bias (offset when calibrated), mse_har_raw (uncalibrated leg) + and mse_har_debiased (calibrated leg) must be finite AND distinct when + debias=True. The c.953 fake identity came from both legs returning the + same dummy forecasts.""" + from har_lj_asym import _eval_one_coin + + def fake_walk_forward_har(rv, horizon, *args, **kwargs): + n = 100 + idx = rv.index[-n:] + offset = -0.05 if kwargs.get("calibrate_bias") else 0.0 + return { + "forecasts": pd.Series( + np.log(rv.iloc[-n:].to_numpy()) + offset, + index=idx, name="fc_har", + ), + "aggregate_mse_logrv": 0.9, + } + + monkeypatch.setattr( + "har_lj_asym.walk_forward_har", fake_walk_forward_har, + ) + + row = _eval_one_coin( + "BTC-USD", horizon=1, seed=0, components=synthetic_lj_components, + debias=True, calibration_size=60, + ) + assert row is not None + assert not np.isnan(row["mse_har_raw"]) + assert not np.isnan(row["mse_har_debiased"]) + assert row["mse_har_raw"] != row["mse_har_debiased"] + + +# --------------------------------------------------------------------------- +# ROUND-3 concern #4 — full-walk-forward OOS-target invariance +# --------------------------------------------------------------------------- + + +def test_walk_forward_lj_asym_oos_target_invariance(monkeypatch): + """Round-3 concern #4: perturbing ALL OOS targets through the FULL + walk_forward_lj_asym must leave ``per_fold_bias`` and + ``forecasts_debiased`` bit-identical (the bias reads the train tail + only, and the OLS forecasts depend on the model + features, not on the + OOS targets). The perturbation is applied by patching + ``realized_variance_to_log`` — the module-level import used ONLY to + build the target — so the features (built by ``lj_asym_features``, + which logs directly) stay untouched. Conversely, perturbing the train + calibration tail MUST move the bias estimate (sensitivity control).""" + import har_lj_asym as module + from har_lj_asym import lj_asym_features, walk_forward_lj_asym + from realized_variance import realized_variance_to_log + + rng = np.random.default_rng(7) + index = pd.date_range("2020-01-01", periods=400, freq="D") + log_rv = np.cumsum(rng.normal(0.0, 0.1, 400)) - 5.0 + rv = pd.Series(np.exp(log_rv), index=index, name="rv") + rv_neg = pd.Series(rv.to_numpy() * 0.5, index=index) + rv_pos = pd.Series(rv.to_numpy() * 0.5, index=index) + rv_c = pd.Series(rv.to_numpy() * 0.8, index=index) + rv_j = pd.Series(rv.to_numpy() * 0.2, index=index) + + horizon, n_splits, calibration_size = 1, 1, 60 + kwargs = dict( + horizon=horizon, seed=0, n_splits=n_splits, + debias=True, calibration_size=calibration_size, + ) + + out_base = walk_forward_lj_asym( + rv, rv_neg, rv_pos, rv_c, rv_j, **kwargs, + ) + assert out_base["forecasts"] + + # Replicate the internal split arithmetic to locate, in ORIGINAL rv + # coordinates, the boundary between train-read and OOS-read targets: + # the target at merged row j reads original position + # (first_pos + j + horizon). + feat = lj_asym_features(rv_neg, rv_pos, rv_c, rv_j, rv) + merged = feat.join( + realized_variance_to_log(rv).rename("log_rv"), how="inner", + ).dropna() + n_total = int( + merged["log_rv"].rolling(horizon).mean().shift(-horizon).notna().sum() + ) + fold_size = n_total // (n_splits + 1) + first_pos = int(rv.index.get_loc(merged.index[0])) + cutoff = first_pos + fold_size + horizon # OOS targets read >= cutoff + + delta = 10.0 + orig = module.realized_variance_to_log + + def shift_from(cut_lo, cut_hi): + def patched(series): + out = orig(series).copy() + mask = (np.arange(len(out)) >= cut_lo) & ( + np.arange(len(out)) < cut_hi + ) + out.iloc[mask] = out.iloc[mask] + delta + return out + return patched + + # (a) Shift ALL OOS targets (original positions >= cutoff): features + # and the train fold are untouched -> bias + debiased forecasts must + # be identical to the baseline run. + monkeypatch.setattr( + module, "realized_variance_to_log", shift_from(cutoff, len(rv)), + ) + out_oos_shifted = walk_forward_lj_asym( + rv, rv_neg, rv_pos, rv_c, rv_j, **kwargs, + ) + # Sanity on the perturbation itself: the OOS targets really changed. + assert not np.allclose(out_oos_shifted["targets"], out_base["targets"]) + np.testing.assert_allclose( + out_oos_shifted["per_fold_bias"], out_base["per_fold_bias"], + rtol=1e-12, atol=1e-12, + ) + np.testing.assert_allclose( + out_oos_shifted["forecasts"], out_base["forecasts"], rtol=1e-12, + ) + np.testing.assert_allclose( + out_oos_shifted["forecasts_debiased"], + out_base["forecasts_debiased"], rtol=1e-12, + ) + + # (b) Shift the train calibration tail instead: the bias estimate MUST + # move (train sensitivity is the expected behavior of a train-only + # estimator). + monkeypatch.setattr( + module, "realized_variance_to_log", + shift_from(cutoff - calibration_size, cutoff), + ) + out_tail_shifted = walk_forward_lj_asym( + rv, rv_neg, rv_pos, rv_c, rv_j, **kwargs, + ) + assert abs( + out_tail_shifted["per_fold_bias"][0] - out_base["per_fold_bias"][0] + ) > 1.0 + + +# --------------------------------------------------------------------------- +# ROUND-3 concern #6 — panel_hash covers the index in addition to the values +# --------------------------------------------------------------------------- + + +def test_panel_hash_includes_index(): + """Round-3 concern #6: two panels with identical VALUES but different + indexes must hash differently (the index bytes participate in the + digest); identical panels hash identically; different values on the + same index hash differently.""" + from har_lj_asym import _panel_hash + + rng = np.random.default_rng(3) + vals = np.exp(rng.normal(-5.0, 0.5, 100)) + idx_a = pd.date_range("2020-01-01", periods=100, freq="D") + idx_b = pd.date_range("2021-03-01", periods=100, freq="D") + + h_a = _panel_hash(pd.Series(vals, index=idx_a)) + h_a_again = _panel_hash(pd.Series(vals, index=idx_a)) + h_b = _panel_hash(pd.Series(vals, index=idx_b)) + h_a_shifted_vals = _panel_hash(pd.Series(vals * 1.01, index=idx_a)) + + assert h_a == h_a_again + assert h_a != h_b, "index must participate in the panel hash" + assert h_a != h_a_shifted_vals, "values must participate in the hash" + + +# --------------------------------------------------------------------------- +# ROUND-5 concern (2) — _content_hash(index, bornes, values) discriminates +# any mutation of one of the three components. +# --------------------------------------------------------------------------- + + +def test_content_hash_mutates_on_index_shift(): + """Round-5 concern (2): shifting the index by one period must change the + digest (even when bornes and values stay identical).""" + from har_lj_asym import _content_hash + + rng = np.random.default_rng(7) + values = rng.normal(0.0, 1.0, 30) + idx_a = np.arange(1_000_000, 1_000_030, dtype=np.int64) + idx_b = np.arange(1_000_001, 1_000_031, dtype=np.int64) + bornes = (10, 10, 30) + + h_a = _content_hash(idx_a, bornes, values) + h_b = _content_hash(idx_b, bornes, values) + assert h_a != h_b, "index shift must mutate the content hash" + + +def test_content_hash_mutates_on_bounds_change(): + """Round-5 concern (2): changing one of the bornes (train_end, oos_start, + oos_end) must change the digest (even when index and values stay).""" + from har_lj_asym import _content_hash + + rng = np.random.default_rng(11) + values = rng.normal(0.0, 1.0, 30) + idx = np.arange(2_000_000, 2_000_030, dtype=np.int64) + bornes_a = (10, 10, 30) + bornes_b = (11, 11, 30) # train_end_idx +1 + + h_a = _content_hash(idx, bornes_a, values) + h_b = _content_hash(idx, bornes_b, values) + assert h_a != h_b, "bounds change must mutate the content hash" + + +def test_content_hash_mutates_on_value_change(): + """Round-5 concern (2): changing one value byte must change the digest. + Sanity guard against a future refactor that drops the values component + (e.g. caching only index + bornes).""" + from har_lj_asym import _content_hash + + rng = np.random.default_rng(13) + values_a = rng.normal(0.0, 1.0, 30) + values_b = values_a.copy() + values_b[15] += 1e-9 + idx = np.arange(3_000_000, 3_000_030, dtype=np.int64) + bornes = (10, 10, 30) + + h_a = _content_hash(idx, bornes, values_a) + h_b = _content_hash(idx, bornes, values_b) + assert h_a != h_b, "value change must mutate the content hash" + + +def test_content_hash_is_deterministic(): + """Round-5 concern (2): identical inputs (index, bornes, values) must + hash identically across repeated calls (deterministic, no hidden state).""" + from har_lj_asym import _content_hash + + rng = np.random.default_rng(17) + values = rng.normal(0.0, 1.0, 30) + idx = np.arange(4_000_000, 4_000_030, dtype=np.int64) + bornes = (10, 10, 30) + + h1 = _content_hash(idx, bornes, values) + h2 = _content_hash(idx, bornes, values) + assert h1 == h2 + + +# --------------------------------------------------------------------------- +# ROUND-5 concern (1) — per_fold_bounds aligned 1-pour-1 with per_fold_bias +# --------------------------------------------------------------------------- + + +def test_per_fold_bounds_aligned_with_per_fold_bias( + synthetic_lj_components, monkeypatch, +): + """Round-5 concern (1): per_fold_bounds must list one entry per fold, + aligned 1-pour-1 with per_fold_bias. Each entry must carry + {fold_idx, train_end_idx, oos_start_idx, oos_end_idx, n_train, n_oos} + and the (train_end_idx, oos_start_idx) must match the walk-forward + geometry: split = (fold_idx + 1) * fold_size.""" + from har_lj_asym import ( + walk_forward_lj_asym, + ) + components = synthetic_lj_components + # Drive walk_forward_lj_asym via the public surface: 3 folds, BTC-USD. + n_splits = 3 + res = walk_forward_lj_asym( + rv=components["BTC-USD"]["rv"], + rv_neg=components["BTC-USD"]["rv_neg"], + rv_pos=components["BTC-USD"]["rv_pos"], + rv_c=components["BTC-USD"]["rv_c"], + rv_j=components["BTC-USD"]["rv_j"], + horizon=1, + seed=0, + n_splits=n_splits, + ) + biases = res["per_fold_bias"] + bounds = res["per_fold_bounds"] + # fold_size = n // (n_splits + 1), where n = len(X_all). Recover it from + # bounds_train_test (the arithmetic is canonical from round-4). + bt = res["bounds_train_test"] + assert bt is not None, "bounds_train_test missing -- needed to derive fold_size" + fold_size = bt["fold_size"] + assert len(biases) == len(bounds), ( + f"per_fold_bounds length {len(bounds)} must equal per_fold_bias " + f"length {len(biases)}" + ) + for k, (b, bd) in enumerate(zip(biases, bounds)): + assert bd["fold_idx"] == k + expected_split = (k + 1) * fold_size + assert bd["train_end_idx"] == expected_split + assert bd["oos_start_idx"] == expected_split + assert bd["oos_end_idx"] == expected_split + fold_size + assert bd["n_train"] == expected_split + assert bd["n_oos"] == fold_size + + +# --------------------------------------------------------------------------- +# ROUND-4 concern (a) — multi-fold OOS-target invariance (n_splits=3, +# distinct per-fold biases, per-fold assertions) +# --------------------------------------------------------------------------- + + +def test_walk_forward_lj_asym_oos_target_invariance_multi_fold(monkeypatch): + """Round-4 concern (a): with ``n_splits=3`` (a single fold cannot + distinguish a per-fold bias estimate from a global mean — there is no + second fold to compare against), the walk-forward must produce + DISTINCT per-fold bias estimates, and each fold's bias + forecasts must + be bit-identical under a perturbation of THAT fold's OOS targets only. + + Geometry (X_all/merged-valid coordinates): fold k trains on + ``[0, (k+1)*fold_size)`` and tests on ``[(k+1)*fold_size, + (k+2)*fold_size)``. In ORIGINAL rv coordinates the targets of fold k's + TEST rows read positions ``[first_pos + (k+1)*fold_size + horizon, + first_pos + (k+2)*fold_size + horizon)`` — patched via + ``realized_variance_to_log`` (target-only import; features never go + through it). + + Assertions per fold k (OOS window of fold k shifted by +10): + - ``per_fold_bias[k]`` + fold k's ``forecasts``/``forecasts_debiased`` + slices are bit-identical (rtol 1e-12) — the bias reads the train tail + only, and the OLS fit never sees the OOS targets; + - folds BEFORE k are fully unchanged (their train and test never read + positions >= cutoff_k) — the backward "no cross-fold leakage" + direction; + - fold k's ``targets`` slice really changed (the perturbation landed); + - folds AFTER k are EXPECTED to differ: fold k's test block IS part of + their train (walk-forward expanding window — future-as-train by + design, not leakage). The test asserts their forecast slices differ, + proving the perturbation reached their estimator through the train + path. + + Train-tail sensitivity per fold: shifting ONLY fold k's calibration + tail (``[cutoff_k - calibration_size, cutoff_k)``) moves + ``per_fold_bias[k]`` by more than 1.0 (delta=10 on the synthetic + scale) while earlier folds stay unchanged. + """ + import har_lj_asym as module + from har_lj_asym import lj_asym_features, walk_forward_lj_asym + from realized_variance import realized_variance_to_log + + rng = np.random.default_rng(11) + index = pd.date_range("2020-01-01", periods=400, freq="D") + log_rv = np.cumsum(rng.normal(0.0, 0.1, 400)) - 5.0 + rv = pd.Series(np.exp(log_rv), index=index, name="rv") + rv_neg = pd.Series(rv.to_numpy() * 0.5, index=index) + rv_pos = pd.Series(rv.to_numpy() * 0.5, index=index) + rv_c = pd.Series(rv.to_numpy() * 0.8, index=index) + rv_j = pd.Series(rv.to_numpy() * 0.2, index=index) + + horizon, n_splits, calibration_size = 1, 3, 60 + kwargs = dict( + horizon=horizon, seed=0, n_splits=n_splits, + debias=True, calibration_size=calibration_size, + ) + + out_base = walk_forward_lj_asym(rv, rv_neg, rv_pos, rv_c, rv_j, **kwargs) + assert out_base["forecasts"] + biases = out_base["per_fold_bias"] + + # (1) Multi-fold: >= 3 per-fold bias entries, not 1. + assert len(biases) == n_splits, ( + f"expected {n_splits} per-fold biases, got {len(biases)}" + ) + # (2) DISTINCT bias per fold — the discriminator against a global-mean + # collapse (a single shared bias would make len(set) == 1). + print(f"per_fold_bias (baseline, n_splits={n_splits}): {biases}") + assert len(set(biases)) >= 2, ( + f"per-fold biases collapsed to a single value: {biases} -- " + "a global-mean estimator would be indistinguishable from this" + ) + + # Replicate the internal split arithmetic to locate, in ORIGINAL rv + # coordinates, the OOS target windows: the target at X_all row j reads + # original position (first_pos + j + horizon). + feat = lj_asym_features(rv_neg, rv_pos, rv_c, rv_j, rv) + merged = feat.join( + realized_variance_to_log(rv).rename("log_rv"), how="inner", + ).dropna() + n_total = int( + merged["log_rv"].rolling(horizon).mean().shift(-horizon).notna().sum() + ) + fold_size = n_total // (n_splits + 1) + first_pos = int(rv.index.get_loc(merged.index[0])) + cutoffs = [ + first_pos + (k + 1) * fold_size + horizon for k in range(n_splits) + ] + # Sanity on the geometry replication: forecasts are whole folds. + assert len(out_base["forecasts"]) == n_splits * fold_size + + fs = fold_size + fc_base = np.asarray(out_base["forecasts"], dtype=float) + fcd_base = np.asarray(out_base["forecasts_debiased"], dtype=float) + tgt_base = np.asarray(out_base["targets"], dtype=float) + + delta = 10.0 + orig = module.realized_variance_to_log + + def shift_from(cut_lo, cut_hi): + def patched(series): + out = orig(series).copy() + mask = (np.arange(len(out)) >= cut_lo) & ( + np.arange(len(out)) < cut_hi + ) + out.iloc[mask] = out.iloc[mask] + delta + return out + return patched + + # --- Per-fold OOS invariance --- + for k in range(n_splits): + lo = cutoffs[k] + hi = cutoffs[k + 1] if k + 1 < n_splits else len(rv) + monkeypatch.setattr( + module, "realized_variance_to_log", shift_from(lo, hi), + ) + out_k = walk_forward_lj_asym(rv, rv_neg, rv_pos, rv_c, rv_j, **kwargs) + sl = slice(k * fs, (k + 1) * fs) + + # Anti-leak: fold k's bias and forecast slices bit-identical. + np.testing.assert_allclose( + out_k["per_fold_bias"][k], biases[k], rtol=1e-12, atol=1e-12, + err_msg=f"fold {k}: bias must not read fold {k}'s OOS targets", + ) + np.testing.assert_allclose( + np.asarray(out_k["forecasts"])[sl], fc_base[sl], rtol=1e-12, + err_msg=f"fold {k}: raw forecasts must be OOS-target invariant", + ) + np.testing.assert_allclose( + np.asarray(out_k["forecasts_debiased"])[sl], fcd_base[sl], + rtol=1e-12, + err_msg=f"fold {k}: debiased forecasts must be OOS-target invariant", + ) + + # Backward folds: fully unchanged (no backward cross-fold leakage). + for j in range(k): + np.testing.assert_allclose( + out_k["per_fold_bias"][j], biases[j], rtol=1e-12, atol=1e-12, + err_msg=f"backward leakage: fold {j} bias moved under fold {k} shift", + ) + np.testing.assert_allclose( + np.asarray(out_k["forecasts_debiased"])[j * fs:(j + 1) * fs], + fcd_base[j * fs:(j + 1) * fs], rtol=1e-12, + err_msg=f"backward leakage: fold {j} forecasts moved under fold {k} shift", + ) + + # Sanity on the perturbation itself: fold k's targets really changed. + assert not np.allclose( + np.asarray(out_k["targets"])[sl], tgt_base[sl], + ), f"fold {k}: OOS perturbation did not land on the targets" + + # Forward folds: fold k's test block is part of their TRAIN + # (expanding window) — they MUST NOT be bit-identical. This is the + # train path, not an OOS read: their forecasts absorb the shift + # through the refit, which proves the perturbation was real. + for j in range(k + 1, n_splits): + assert not np.allclose( + np.asarray(out_k["forecasts"])[j * fs:(j + 1) * fs], + fc_base[j * fs:(j + 1) * fs], + ), ( + f"forward fold {j}: expected a train-path change under fold " + f"{k}'s OOS shift (expanding window retrains on fold {k}'s " + "test block) -- bit-identity here would mean the walk-forward " + "silently ignored train data" + ) + + # --- Per-fold train-tail sensitivity --- + for k in range(n_splits): + lo = cutoffs[k] - calibration_size + hi = cutoffs[k] + monkeypatch.setattr( + module, "realized_variance_to_log", shift_from(lo, hi), + ) + out_t = walk_forward_lj_asym(rv, rv_neg, rv_pos, rv_c, rv_j, **kwargs) + + # Fold k's bias moves by a non-trivial margin on the synthetic scale. + d_bias_k = out_t["per_fold_bias"][k] - biases[k] + assert abs(d_bias_k) > 1.0, ( + f"fold {k}: calibration-tail shift ({delta=}) moved the bias by " + f"only {d_bias_k:+.6f} -- the estimator looks insensitive to its " + "own train tail" + ) + # Earlier folds' biases are unchanged (tail window beyond their train). + for j in range(k): + np.testing.assert_allclose( + out_t["per_fold_bias"][j], biases[j], rtol=1e-12, atol=1e-12, + err_msg=( + f"fold {j} bias moved under fold {k}'s calibration-tail " + "shift -- the tail window leaked backwards" + ), + ) + + +# --------------------------------------------------------------------------- +# ROUND-4 concern (c) — bounds provenance (train/OOS index bounds surfaced +# end-to-end: walk_forward_lj_asym -> _eval_one_coin -> aggregate_verdicts) +# --------------------------------------------------------------------------- + + +def test_bounds_provenance_in_manifest(synthetic_lj_components, monkeypatch): + """Round-4 concern (c): the per-(coin, horizon, seed) result must carry + ``bounds_train_test`` matching the walk-forward split arithmetic — + ``train_end_idx = n_splits * (n // (n_splits + 1))``, + ``oos_start_idx = train_end + horizon``, ``oos_end_idx = n`` — plus the + per-fold granules (``per_fold_bias`` + ``fc_lj_hash_per_fold``) that + anchor the global hashes to the bounds. ``aggregate_verdicts`` relays + the bounds per (coin, horizon) with a cross-seed consistency flag (the + same pattern the manifest writes as ``bounds_per_coin_horizon``).""" + from har_lj_asym import N_SPLITS, _eval_one_coin, lj_asym_features + from realized_variance import realized_variance_to_log + + def fake_walk_forward_har(rv, horizon, *args, **kwargs): + n = 100 + idx = rv.index[-n:] + return { + "forecasts": pd.Series(np.zeros(n), index=idx, name="fc_har"), + "aggregate_mse_logrv": 0.9, + } + + monkeypatch.setattr("har_lj_asym.walk_forward_har", fake_walk_forward_har) + + horizon = 1 + row = _eval_one_coin( + "BTC-USD", horizon=horizon, seed=0, + components=synthetic_lj_components, + debias=True, calibration_size=60, + ) + assert row is not None + + # Replicate the LJ walk-forward geometry (same arithmetic as + # walk_forward_lj_asym, which _eval_one_coin calls with the default + # n_splits=N_SPLITS). + comp = synthetic_lj_components["BTC-USD"] + feat = lj_asym_features( + comp["rv_neg"], comp["rv_pos"], comp["rv_c"], comp["rv_j"], comp["rv"], + ) + merged = feat.join( + realized_variance_to_log(comp["rv"]).rename("log_rv"), how="inner", + ).dropna() + n_total = int( + merged["log_rv"].rolling(horizon).mean().shift(-horizon).notna().sum() + ) + fold_size = n_total // (N_SPLITS + 1) + expected_train_end = N_SPLITS * fold_size + + b = row["bounds_train_test"] + assert b is not None, "bounds_train_test missing from the result row" + assert b["train_end_idx"] == expected_train_end == N_SPLITS * (n_total // (N_SPLITS + 1)) + assert b["oos_start_idx"] == expected_train_end + horizon + assert b["oos_end_idx"] == n_total + assert b["n_train"] == expected_train_end + assert b["n_oos"] == n_total - expected_train_end + + # Manifest write path: the bounds must be JSON-serializable as-is. + json.dumps(b) + + # Per-fold provenance granules aligned with per_fold_bias. + assert len(row["per_fold_bias"]) == N_SPLITS + assert len(row["fc_lj_hash_per_fold"]) == N_SPLITS + for h16 in row["fc_lj_hash_per_fold"]: + assert isinstance(h16, str) and len(h16) == 16 + int(h16, 16) # hex parseable + + # aggregate_verdicts relays the bounds per (coin, horizon) with the + # cross-seed consistency flag (deterministic OLS -> identical bounds). + rows = [] + for seed in (0, 7): + r = _eval_one_coin( + "BTC-USD", horizon=horizon, seed=seed, + components=synthetic_lj_components, + debias=True, calibration_size=60, + ) + assert r is not None + rows.append(r) + agg = aggregate_verdicts(rows)[0] + assert agg["bounds_train_test"] == rows[0]["bounds_train_test"] + assert agg["bounds_consistent_across_seeds"] is True