Repository navigation
test(replay): replay a Predbat log forwards from a debug yaml - exact and simulated, with a plan timeline - #5363
Draft
chalfontchubby wants to merge 32 commits into
Draft
chalfontchubby wants to merge 32 commits into
chalfontchubby wants to merge 32 commits into
Conversation
A debug yaml is one moment; bug reports usually also carry the following hours of log. The new --replay_log mode restores the yaml, then steps through the log run by run, setting the clock, SoC, load/PV/import/export history and the inverter's export window from the log, re-planning on the runs where the log re-planned, and printing the logged export windows next to the replayed ones. run_single_debug's state restore and load/PV model rebuild are extracted into helpers so both modes share them. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…plan The candidate charge/export windows start at the current slot, and fetch rebuilds them every cycle. The replay never did, so every re-plan optimised the yaml's original windows and the plan froze. Extract the --redo rate rescan into rescan_rate_windows() and call it before each replayed re-plan. On a 3 Oct Sigenergy log replayed from the 07:00 yaml the first export window's start now matches the log on 12 of 15 re-plans (9 before). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…xport targets --replay_chart PNG draws the log's actual SoC with the export target of the live plan and the replayed plan on the same % scale, plus a lane per plan showing whether it says export (or freeze) at each run. Uses the Agg backend, so it never opens a window. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two modes now: - exact (default): every re-plan starts from the SoC the log recorded, to check the replay reproduces the live plans; - simulated (--replay_simulate): the battery is stepped forward under the replayed plan with the actual PV and load from the log, using Predbat's own battery and inverter model (rate curves, losses, reserve, inverter and export limits). This is the mode for trying a code change, since the SoC then follows the changed plan. Plan windows are now parsed with their date, as minutes from the run's midnight. Previously tomorrow's 06:00 window read as today's, so the chart showed exports that were never planned for today. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
calculate_plan recomputes the load divergence from the load history, and the replay can only rebuild that history at the log's 5-minute resolution, so the divergence came out smoother than the live system's. The replay now overrides get_load_divergence with the value logged for each run (restored afterwards). The cloud factor is no longer injected: calculate_plan derives it from the forecasts, which the replay already has. Issue 4865 from 05:30: 22 of 51 re-plans now match the logged first export window start (20 before). Rik's 3 Oct log from 04:00: 8 of 34 plans identical (5 before). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y SoC RMS error A dummy_inverter apps.yaml block enables a component that publishes inverter sensors and controls and wires Predbat to them, so Predbat plans and controls it like a real inverter - for demos, and as an independent inverter model for log replay. Behind the entities is a minute-by-minute hybrid model: rate limits and losses, PV and battery sharing the inverter's AC limit (PV first), export capped at the export limit, clipped PV counted, charge/export windows obeyed (zero rate = freeze), Eco otherwise. A DUMMY row in INVERTER_DEF carries its capabilities. The replay now also prints the RMS difference between simulated and logged SoC: 3.18% on the #4865 day, 0.50% on Rik's 3 Oct log. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…uilds each run Fidelity fixes found by stepping through the first divergences of the #4865 replay: - compare the adopted plan with the log's adopted plan (the next run's first "Best export window"), and the candidate plan with the log's candidate ("Export windows filtered") - before, the log's candidate was compared with the replay's adopted plan, which differ whenever a re-plan is rejected; - rebuild the weighted-bucket historical load forecast before each re-plan, as fetch does; - age the car charging and iBoost histories with the load history, since the load filter subtracts them - unshifted, overnight car charging was misaligned and the load forecast drifted upwards; - take the cost so far today and the day counters from the log, as every metric starts from them; - set the inverter's export limits alongside its export window, and publish while re-planning so the replay logs the same "predict base/best" lines as the live system for per-run comparison. Rik's 3 Oct log from 04:00: 20 of 34 adopted plans now identical to the log (8 before). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A faithful replay needs the code that wrote the log. The log records the running version every run, so the replay now stops at the first run whose version differs and says so - the #4865 log, for one, spans an upgrade from v9.3.1 to v9.3.3. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rows from the first run on a new version are flagged, the summary counts them separately, and the chart marks the change. A replay past it is still useful but not expected to match as closely. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Predbat versions with feat/log-replay-inputs write "Replay input:" lines for the PV forecast (on each change) and the load forecast (each cycle). The replay now applies them: a logged PV forecast replaces the yaml's from its half hour onwards, and a logged load forecast replaces the rebuilt one before the re-plan. Logs without the lines replay as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… replays Applied after the debug yaml is restored, repeatable, with values read as numbers or booleans where they look like them; an unknown setting name is rejected. First use: replaying the #4865 morning in simulated mode with pv_metric90_weight at 0.15, 0.25 and 0.5. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e debug replays too What-if overrides are useful for a single debug yaml as much as for a forward replay, so the option is now --override name=value and run_single_debug applies it after restoring the yaml. The shared helper apply_overrides lives beside restore_debug_state. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…put lines, the 8-hourly config refresh re-plan Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…w on every replayed re-plan Fetch recomputes them each run over the horizon from now, so a passed peak drops out. The replay kept the yaml's values, which after midnight would still include yesterday's rates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the one logged after executing it A run logs the inverter's SoC before planning and again after executing the plan; the parser kept the last, so every replayed re-plan started a few minutes' battery movement away from the live one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
At midnight the replay moves everything held as minutes from midnight (rates, PV and load forecasts, keeps, plan and inverter windows, manual times) back a day and advances midnight_utc; plan_last_updated_minutes is left alone so calculate_plan forces the start-of-day re-plan as the live system does. Day counters that restart at midnight count from zero, the midnight run's tariff comparison plans are ignored, and an export window over midnight seen in the small hours started yesterday. --replay_until HH:MM means the first such time after the yaml. Replaying the 23:00 yaml of 3 Oct through to 09:50 on 4 Oct: 28 of 65 adopted plans identical to the log, 49 with the same first export start; simulated SoC within 2.6% RMS of the log over 131 runs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When the 8-hourly config refresh is pending the run's adopted plan is marked invalid, and the next run (at the same minutes_now) re-plans and adopts without comparing to the old plan. The replay now does the same, and no longer reads that run's working window list as the plan in force. Without it the replay kept the old plan under metric_min_improvement_plan and stayed on it for hours. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…es; mark each replayed run in the replay's log fetch_sensor_data resets pv_forecast_minute90_signatures every cycle, so the plan's stale-p90 guard only compares within one cycle. The replay never cleared them, so a logged PV update that moved p50 but left p90's values alone read as a p90 left behind: the plan swapped p90 for p50, dropped to the proportional cloud model and ran 5-7p off the live metric from then on. On the 3-4 Oct log this takes adopted plans identical to the log from 29 to 55 of 71. The replay's own predbat.log now carries a line per replayed run, so it can be read beside the live log. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Predbat now logs "Replay input: state" every cycle (SoC, soc_max, in-day adjustment, cost so far and the day counters, each as a repr that reads back as the same float) and "Replay input: rates changed" whenever the import/export rates or their base curves change. The replay now: - takes the plan's starting values from the state line over the rounded human-readable lines, and uses its exact counters for the history shifts and the simulated battery's PV and load; - rebuilds the four rate series from the logged change points, so rates fetched after the yaml (the next day's Agile prices, a new futurerate prediction) reach the re-plans. Rates are logged only on change, so a line from a run the parser drops carries to the next run, moved back a day if that run is past midnight. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"Will recompute the plan as it is invalid" is logged for every recompute, whatever triggered it, so the replay treated each sensor-triggered re-plan as invalid and adopted the new plan without the old-vs-new comparison the live run made. Only calculate_plan's "Recompute, previous plan is invalid..." marks a really invalid plan; the first line now only stops a recomputing run's fresh window list being read as the plan in force. Replay of the 11:30 yaml to 16:50 on 4 Oct: adopted plans 26 of 28 identical (22 before), candidates 31 of 31. Also corrects the journal's 8-hourly refresh bullet, which still said the replay could not follow that re-plan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Newer logs carry, for every step the plan builds, the two cumulative load forecast values step_data_history() reads (the step's start and the minute after), exactly. The replay sets those and leaves other minutes alone; older logs' per-slot line is still read when the exact one is absent. Checked by injecting the line, built from the live 14:50 debug yaml, into the 4 Oct log: the replay's load_minutes_step (and 10/90) then match live's at 14:50 exactly, as do predict_metric_best and predict_soc_best. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Newer logs give p50/p10/p90 per minute from now to the end of the plan, exactly, as runs of [kWh, minutes]; the replay sets those minutes (removing ones logged None) and leaves the rest. The older half-hourly line is still read when the exact one is absent. Checked by injecting the exact load and PV lines, built from the live 14:50 debug yaml, into the 4 Oct log: the replay's 14:50 re-plan then matches live's line for line, the previous plan's re-score included, and its load/PV step series, predict_metric_best and predict_soc_best equal live's exactly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…only on re-plans The replay applied the logged load forecast only on runs that re-planned, so between re-plans the instance held the previous re-plan's forecast while live held a fresh one. No re-plan changes, but a state dump at a non-re-plan run (the hourly debug yamls fall on :15, re-plans on :x0) now compares like for like. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"Replay input: cars changed" carries the run-time car fields the plan reads, logged only on change; the replay sets them each run it applies, and carries a line from a dropped run on to the next kept run, as it does the rates. Checked on the 5 Oct log by injecting the car state from the live 18:15 debug yaml at 17:25, when the car was plugged in: adopted plans 47 of 48 identical (43 of 46 without it), and the 18:15 re-plan, where the injected state is the real one, matches live line for line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the exact load divergence Where a log carries "Replay input: inverter changed", the replay sets the inverter's programmed state from it and keeps it between changes, instead of rebuilding the export window from the previous run's lines. The rates in force come from the state line, and the exact divergence wins over the rounded percentage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mperature from the state line The compact line counts each value in whole tenths of a Wh and divides by 10000 only at the end, which gives exactly the dp4 float the live forecast held: the 268 forecasts in an overnight log (195,640 values) all rebuild bit for bit. Logs with the full cumulative kWh line still replay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rt and car - for live and replay The chart's bottom strip showed only export and freeze. It now gives each plan's state at every run in the web plan's terms and colours (Chrg, HoldChrg, FrzChrg, Exp, HoldExp, FrzExp, car), from both the export and the charge windows in force, plus a lane for the car's planned charging. The live charge windows are the "Best charge window" line at the start of the next run, the replay's its own adopted list. plan_state_now() replaces export_mode_now(). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… dynamic_load() does live Live dynamic_load() rebuilds car_charging_now_slots every cycle, so the plan holds the battery for a car that reports charging with no planned slot covering it. The replay never did, so at 00:05 on 6 Oct it planned a battery charge 00:00-00:30 that the live plan did not have. Export windows still matched, which hid it, but a simulated replay charged the battery there and drifted: 8.8% SoC RMS overnight, 1.99% with this. The replay now derives the slots with the live dynamic_load_car_charging_now() from the logged car state. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…plan timeline chart Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…logs it when divergence is off The logging now writes the divergence exactly as the plan uses it, which is None when metric_load_divergence_enable is off. The replay applied the rounded percentage line instead, giving the plan a divergence live never used. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
chalfontchubby
added a commit
that referenced
this pull request
Oct 7, 2026
… a log can be replayed exactly A debug yaml captures one moment; a forward replay of the log after it (draft PR #5363) needs the values each later plan started from, which the human-readable lines round or omit. Each line is prefixed "Replay input:", reads back exactly (repr floats, or dicts that ast.literal_eval reads), and is logged on change where the value rarely changes. None of it changes what Predbat does: every logger catches its own errors, and a value that cannot be logged warns once per change. Tariff comparison runs (save=False / publish=False) log none of it, as their inputs are not the live plan's. - PV forecast, on change: p50/p10/p90 per minute from now to the end of the forecast, as runs of [kWh, minutes]. - Load forecast, every cycle: the two values per 5-minute step the plan reads (step start and the minute after). Rounded to 0.1 Wh, so logged compactly as Wh per step from a cumulative base, about half the size of the full cumulative kWh line it falls back to for a forecast that is not in whole tenths of a Wh. - Load ML predictions, on change, exactly as read from the sensor. - Rates, on change: import/export and their base rates from now, as change points. - Plan starting state every cycle: SoC, SoC max, in-day adjustment, cost so far, the day counters, the charge and discharge rates in force and the battery temperature. - Car state and the inverter's programmed state (windows, charging/exporting and targets, reserve, rate limits), each on change; and the load divergence exactly as the plan uses it (None when divergence is off). The on-change lines share one helper (log_replay_changed). On a live system the replay lines are about 5% of the log. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
chalfontchubby
force-pushed
the
tools/debug-replay-forward
branch
from
October 7, 2026 12:44
92a60d7 to
07a1416
Compare
2 tasks done
…oximately, back to v8.8
A log from before the replay-input logging could not be replayed: the run's SoC, day totals and re-plan marker
had all changed format over the versions, so most older logs gave no runs at all. The replay now reads:
- the SoC from "Inverter 0 SOC: 1.9kW ..." / "SoC: 6.85kW ..." as well as today's line, and the all-inverter
total from "Found N inverters totals" where a log has it (so a multi-inverter system starts from the total);
- the older spaced day totals ("load 5.04 kWh import 14.07 kWh ... pv 0.0 kWh");
- a re-plan from "Filtered charge windows", for versions that do not log the filtered export windows;
- where a run has no exact cars line, the car's planned slots from "Car N charging plan is" (since v5.1) and its
planned and charging-now flags (since v7.0);
- where the log has no rates lines, the Intelligent dispatch list from "Octopus slots changed" (since v8.27.27),
with the import rates rebuilt from it by the live rate_add_io_slots.
A log without any replay-input lines gets a note that the plans are approximate, and one with no readable runs a
clear error. On the 5-6 Oct overnight log with every replay-input line stripped, adopted plans identical rose from
5 to 20 of 66 (same first export start 14 to 66) and the simulated SoC error fell from 11.2% to 2.0%; real issue logs
from v8.8.13 to v9.0.3 now parse, and a 20-hour v9.0.3 log replays. The full-input replay is unchanged (66 of 66).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… line A multi-inverter log logs each inverter's own SoC line, and newer versions dropped the "Found N inverters totals" line, so the replay planned from inverter 0's SoC alone: on the 3-inverter system in #5423 that was 6.46 kWh of a real 18.54 kWh at 18:45, and no re-plan matched. The run's SoC is now the sum of each inverter's first SoC line of the run (matching live's own total, 26.86 vs 26.865 kWh at 17:30), and each inverter is set to its own. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Status: a work in progress across this and later PRs
The replay is a tool we will keep improving over several PRs. It will not reproduce every situation exactly yet, and this PR does not try to close every gap: each case it misses tells us what to log or derive next. Exact replay needs a log with the
Replay input:lines from #5427; older logs replay approximately. Known gaps are listed below and will be taken one at a time.Summary
A debug yaml is one moment in time; the question in a bug report is usually "why did the plan do that at 09:25?" when the yaml is from 05:30. This replays a Predbat log forwards from a debug yaml, one live run at a time, re-planning wherever the live run re-planned, and compares what the replay planned with what the log says live planned.
Two modes
--replay_simulate): the battery is stepped by Predbat's own prediction model using the PV and load that actually happened. This is for trying code or setting changes against real days, and its SoC error against the log measures the battery model itself.--override name=valuechanges a setting for a what-if run.What it reads from the log
Clock, SoC and day counters from the normal log lines, and, where the log has them,
Replay input:lines with the exact values the plan used: rates, PV forecast, load forecast, plan starting state, the car state, the inverter's programmed state and the load divergence. Those lines are added by #5427, which this PR is stacked on. Derived state the live loop rebuilds every cycle (the p90 guard, rate windows, a car charging outside its plan) is rebuilt with the live code rather than logged.It carries across midnight, follows invalid-plan and sensor-triggered re-plans, and marks rows after a logged version change.
Older logs
Logs from before the
Replay input:lines replay too, approximately, back to at least v8.8.13 (Dec 2024). The replay prints a note saying so. It reads the older SoC, day-total and re-plan line formats, and the all-inverter SoC total where a log has it. Where the exact lines are missing it takes the car's planned slots and flags fromCar N charging plan is(since v5.1) and the Cars line (since v7.0). It takes the Intelligent dispatch list fromOctopus slots changed(since v8.27.27), with the import rates rebuilt from it by the live code. The forecasts stay as the yaml had them.Replay input:line strippedReal issue logs from v8.8.13, v8.18.7, v8.25.5, v8.32.14, v8.46.4 and v9.0.3 all parse, and a 20-hour v9.0.3 log (#2716) replays.
Output
--replay_chart: actual SoC (and simulated SoC) against each plan's export target, plus a plan timeline per side in the web plan's terms (Chrg, HoldChrg, FrzChrg, Exp, HoldExp, FrzExp, car) and a lane for the car's planned charging.Replay of run at ...markers so it can be read next to the live log.Fidelity on a live system (Sigenergy, IOG, 6 Oct 2026)
Known gaps
Not in this PR
Changes
apps/predbat/tests/replay_forward.py(new): log parser, replay loop, comparison, simulated mode, chart.apps/predbat/tests/test_replay_forward.py(new): unit tests for parsing, every logged input, history shifting, midnight, the plan timeline and the car modelling.apps/predbat/dummy_inverter.py(new): a simulated inverter and battery component.apps/predbat/tests/test_single_debug.py: helpers shared withrun_single_debug, whose behaviour is unchanged (debug_casesgoldens pass), plus--override.apps/predbat/unit_test.py: the replay flags and test registration.docs/developing.md: how to use the replay.Test plan
./run_all --test replay_forward --test debug_cases./run_all --quickcoverage/run_pre_commit🤖 Generated with Claude Code