From c7c46c50abbd63ac596fb47db8965e7a262047c4 Mon Sep 17 00:00:00 2001 From: CI Date: Mon, 5 Oct 2026 00:16:03 +0100 Subject: [PATCH 1/2] docs(debug-journal): fold in the 2026-10-04 queue slice (9 candidates); mark GH#5376 fixed by PR #5377 (v9.3.5) and GH#5366 by PR #5383; re-verify the SolaX/counter-dip entries after PR #5389; add the CID499 read/write units, marginal-band contamination, GivTCP verify-tolerance, minute-data dip asymmetry, Fox MaxSoc, car-export-blind, re:-first-match and tooltip-clamp entries Co-Authored-By: Claude Code --- tools/debug-journal.md | 23 ++++++++++++----------- 1 file changed, 12 insertions(+), 11 deletions(-) diff --git a/tools/debug-journal.md b/tools/debug-journal.md index 507efb89e..5db3bdbf8 100644 --- a/tools/debug-journal.md +++ b/tools/debug-journal.md @@ -25,7 +25,7 @@ Committed examples of the same format live in `coverage/cases/*.yaml`; they run **A replay cannot reproduce ML-forecast-dependent plan choices.** `--redo` resets the load model, and the live ML forecast cannot be replayed from the dump — so a replay differing from the live plan in exactly the optimiser's marginal ~10p slots is expected, not proof the slot is a bug (GH#5213). Treat "replay doesn't show it" as inconclusive when the candidate choice is worth ~10p; the decisive check is whether the *same version's own* plan in the attached log contains the slot — grep `Best export window`/`Best charge window` lines per version window in their log. The same recipe replays a dump outside `--debug_file`: register a temporary test function in `TEST_REGISTRY` that calls `run_single_debug(..., redo=True)` (`tests/test_single_debug.py`), run via `tools/triage_test.sh` — which is also the only way to exercise the fetch-layer threshold logic (`set_rate_thresholds()` + the `rate_scan_window()` rescan) on a dump's rates, since the dump's stored threshold/windows are live-run end state, useful for the symptom but not for re-derivation (GH#5221). -**A v9.1.0+ dump can crash the loader itself before any Predbat code runs.** Dumps written by v9.1.0+ can carry cached `numpy.dtype` objects (`!!python/object/apply:numpy.dtype`, anchored and aliased through `numpy._core.multiarray.scalar` entries), and the harness's `DEBUG_YAML_LOADER` (`userinterface.py`) raises `SystemError ... bad argument to internal function` from `dtype.__setstate__` on the coverage venv's numpy 2.x while loading (GH#5205 replay, verified three times on a 3.2 MB dump). Lossless replay workaround (throwaway edit next to the loader definition): wrap `yaml.constructor.UnsafeConstructor.set_python_instance_state` and skip state application when `type(instance).__module__` starts with `"numpy"` — the dtype is already fully defined by its constructor args, so skipping the legacy state tuple is lossless. Two guard mistakes that each cost a run: the instance type is not named `dtype` on numpy 2.x (`np.dtype('f8')`'s type is `Float64DType`), and the passthrough must use `*args/**kwargs` — this PyYAML version rejects an `unsafe=` kwarg. If this is ever fixed properly it belongs in the loader definition, not in the replay flow. +**A v9.1.0+ dump can crash the loader itself before any Predbat code runs.** Dumps written by v9.1.0+ can carry cached `numpy.dtype` objects (`!!python/object/apply:numpy.dtype`, anchored and aliased through `numpy._core.multiarray.scalar` entries), and the harness's `DEBUG_YAML_LOADER` (`userinterface.py`) raises `SystemError ... bad argument to internal function` from `dtype.__setstate__` on the coverage venv's numpy 2.x while loading (GH#5205 replay, verified three times on a 3.2 MB dump). Lossless replay workaround (throwaway edit next to the loader definition): wrap `yaml.constructor.UnsafeConstructor.set_python_instance_state` and skip state application when `type(instance).__module__` starts with `"numpy"` — the dtype is already fully defined by its constructor args, so skipping the legacy state tuple is lossless. Two guard mistakes that each cost a run: the instance type is not named `dtype` on numpy 2.x (`np.dtype('f8')`'s type is `Float64DType`), and the passthrough must use `*args/**kwargs` — this PyYAML version rejects an `unsafe=` kwarg. If this is ever fixed properly it belongs in the loader definition, not in the replay flow. The class is wider than numpy: dumps carrying other `python/object` tags (`utils.MinuteArray`, `control_ledger.ControlLedger`) crash `safe_load` the same way — a tolerant loader reconstructs a `MinuteArray` from its serialized `_data` and accepts `python/object/*` generically as attribute dicts (GH#5382 replay, verified 2026-10-04). Two lines in that output are normal and are not the reporter's bug: @@ -102,14 +102,14 @@ Grep for the named symbol rather than trusting a line number. | Area | What past debugging found | Targeted test | |------|---------------------------|---------------| -| Fox (`fox.py`) | The cloud API returns errno 42015/44096 for settings a given device does not support (`FOX_SETTINGS_UNSUPPORTED_ERRNO`); those are marked unavailable and never polled or written again. Entity type matters — WorkMode is a select, ExportLimit a number. Two later capacity/schedule traps, both still live: `publish_data()` sums every `batteryList` entry unconditionally (`fox.py:1779-1781`), and an AIO ESS returns one physical pack as four `bmu` entries all carrying the inverter's own `batterySN`, so `soc_max` comes out at 4x the correct `batteryDesignCapacity` and `soc_kw` with it (GH#4919, read out of the reporter's own API response). `fox_automatic: true` re-`set_arg`s `soc_max` every cycle and `soc_max` is not in `CONFIG_API_OVERRIDE`, so apps.yaml cannot override it - the escape hatch is `fox_automatic: false`. With the Mode Scheduler deleted the scheduler `enable` flag reads 0 and `compute_schedule()` (`fox.py:1080`) derives the displayed charge window from the legacy `forceChargeTime` read, which Predbat cannot clear: `set_battery_charging_time()` (`fox.py:1009`), the only writer of that endpoint, has zero production callers - re-verified on main, where it is referenced only by `test_fox_api.py` (GH#4939). Note this is the *residual* half of that issue: the blocker the reporter actually hit was HA event routing filtering on the literal string `predbat` rather than the configured entity prefix, **fixed in PR #4962**. Unlike the self-referential HA verify trap below, a Fox `didn't complete got X` is a genuine failure - the read-back polls entities the component republishes from Fox Cloud. A separate trap on the **HA/modbus path** (inverter type `FoxESS`, not `FoxCloud`): the stock template binds `reserve:` to `number.foxess_min_soc_on_grid` (`templates/fox.yaml`), and that file's own mode table ties both Freeze charging and Hold charging to that entity. A reporter had `reserve:` bound to `number.foxess_inverter_min_soc` instead - a name that appears nowhere in `templates/` or `docs/` - so Predbat never read or wrote `min_soc_on_grid`; when the inverter stranded it at 100 (observed sitting there six days) Predbat could neither see it nor clear it, and `min_soc_on_grid` at 100% stops the battery discharging while grid-connected. `FoxESS` is `has_idle_time: False` (`config.py`), so the `idle time is ...` line is bookkeeping only and never programs a demand period - discharge outside a charge window depends entirely on the inverter's own settings, which is what makes a wrong reserve binding silent (GH#4961). FoxCloud freeze export (GH#5015, **fixed in PR #5038**, v9.0.2 - keep the mechanism for pre-v9.0.2 logs): `support_feedin_first: True` makes `prediction.py`'s freeze branch model freeze export as genuine Feed-in-First — load exports up to the limit, only the surplus beyond it charges the battery — but the cloud path used to never select the `Feedin` work mode: `apply_battery_schedule()` (`fox.py`) only emitted `SelfUse`/`ForceCharge`/`ForceDischarge` groups and `adjust_inverter_mode()` hardcodes `SelfUse` for Fox (`inverter.py`), so the inverter sat in SelfUse where surplus PV charges the battery *before* exporting and the plan's freeze-export revenue didn't match the hardware. #5038 makes `apply_battery_schedule()` select `Feedin` as the baseline work mode whenever freeze export is requested; note the cloud API's verify read is slow there — a write can log success, read back pre-write 3s later and read back the new mode after ~18s (recorded in a `fox.py` comment), so a quick verify read is not evidence of failure. `support_feedin_first` therefore means two different things depending on connection method: the modelling is correct on the FoxESS HA/modbus path, where `discharge_freeze_service` genuinely selects Feed-in First (#4207, fixed by #4425), while #4582 extended the flag to the cloud defs with no equivalent execution path. Log check: a freeze-export day shows hundreds of full-day `SelfUse` groups in `Fox: New schedule` lines and never a `ForceDischarge` or `Feedin` group — and `ForceDischarge` appearing in a `Fetch scheduler V1 returned` enum list is **not** evidence a discharge window was written. Workaround: turn `set_export_freeze` off so the planner stops selecting freeze-export slots. #4182 is the neighbouring open issue. | `fox_api`, `fox_oauth` | -| Solis (`solis.py`) | `SOLIS_CID_STORAGE_MODE = 636` is Modbus 43110, a bit mask. On firmware "4B and above" (what `is_tou_v2_mode()` detects: CID 6798 reads 43605) the timed charge/discharge enable moved to the per-slot registers — `SOLIS_CID_CHARGE_ENABLE_BASE`/`..._DISCHARGE_ENABLE_BASE`, Modbus 43707, six slots not three — and every mode value carrying TOU bit 1 (3/35/43/51/98) was dropped; 35 became 33, 98 became 96. Such an inverter answers a CID 636 write with code 0 and reads back without bit 1, so a log full of CID 636 verification warnings is a refused bit, not a failed write (GH#4707). `set_storage_mode_if_needed()` decides the bit from `is_tou_v2_mode()`: never asked for on V2, always asked for and retried on V1, where bit 1 still *is* the timed charge/discharge enable. Do not try to learn it from the read-back instead — GH#4710 did, and GH#4774 showed why that cannot work: on six inverters the same write of 179 verifies minutes before and minutes after the one that reads back 177, so a single post-write read is not evidence of a firmware property. Latching it was also silent rather than noisy: the verify read refreshes `cached_values`, so after a stripped write the cache held the stripped value, the suppressed computed value matched it, and `set_storage_mode_if_needed()` stopped writing CID 636 at all for the whole 8-hour verdict — across an overnight charge window. GH#4239's "only retained while a window is configured" is not the whole story — it was refused with slot 1 enabled and its window in force. `read_and_write_cid()` re-reads once after `verify_settle_seconds` before calling a mismatch a failure, because the immediate verify read is taken about half a second after the write. On the SolisCloud select path specifically, **PR #5247 (merged 2026-09-26) makes a discovered inverter's slot-time writes run their handler immediately** so Predbat's read-back sees them in the same cycle — pre-#5247 the select callback only queued the event (GH#4875) and the queue drained at the top of `run()` up to a minute later, so the write verified against the stale value and logged a spurious `didn't complete` once per moved slot time (pass 1, V1 and the slot SoC/current writes keep their every-minute retry). Do not "stop the failed verify poisoning the cache" — `write_cid()` caches the value it *requested* and the post-write `read_cid()` deliberately overwrites it with what the inverter actually reports, so the cache mirrors the inverter rather than Predbat's intent; skipping that update would leave the cache agreeing with the write that just failed, and change detection would then never retry it. The two control paths choose the storage mode differently and it matters: V1 picks from `in_charge_slot`/`in_discharge_slot`, which are clock tests, so CID 636 changes value and gets written the moment a window opens; V2 picks from `slot1_active`, which is only "slot 1 has a window configured", so the mode is written when the slot is *programmed* and nothing at all happens at the window boundary. On the GH#4774 night the last CID 636 write was 22 minutes before the window opened, and the comparison night that worked had one near the window - so `claim_window_mode_assertion()` now asserts the mode once per window on both paths. Slot registers are polled hourly, plus every 5 minutes while `is_inside_active_window()` is true, so drift during the window that matters is caught in the same cycle it happens. GH#4774 also left a useful negative result: that inverter grid-charged from 20% to 91% overnight with bit 1 clear throughout, confirming the per-slot enables alone drive timed charging on V2 — so a "charge window never started" report on V2 firmware is not a CID 636 problem and needs looking at elsewhere. Separately, an inverter reporting `batteryType 'No Battery'` stays in `self.inverter_sn` for polling; `is_battery_inverter()` keeps control writes off it, which was most of the warning volume in GH#4707. That explanation doesn't cover every CID 636 report, though: a later GH#4707 comment showed zero `setting storage mode to` idle-mode log lines, and both V1 mode-decision idle paths always log that line — so a report with neither line cannot be coming from `write_time_windows_if_changed()` at all. The only remaining write path is the HA select handler, `set_storage_mode_value()` at `solis.py:2574`. Check which write path actually ran before assuming this entry's explanation applies. Separately, on inverters using `H M` time format (`GS_fb00`, most cloud subtypes) `adjust_force_export()` used to re-commit a stable export window every cycle — `is_hm_format` alone made `changed_start_end` true, so `press_and_poll_button()` fired on every cycle of an unchanged export window (GH#4709, confirmed live on main at the time) — and PR #4713 fixed the twin idle-cycle press where the times being managed came back as `None` and never compared equal to what the inverter still reported (GH#4712, #2328). **Both merged together in PR #4711 (`bc853a0e`, 2026-09-12, post-v9.0.2)**: `adjust_force_export()` re-commits only on a real schedule change through the commit-once ledger (`last_committed`/`commit_pending`, `commit_needed()`/`record_commit()` in `inverter.py` — the attribute this entry used to name, `last_export_schedule_committed`, is gone), and a commit is only recorded when the writes *and* the button press both succeeded; **the guard only became real when PR #5126 (merged 2026-09-26, unreleased at the time of writing) stopped rebuilding Inverter objects every cycle** — pre-#5126 the rebuild reset the latch to empty each cycle, so the guard was decorative and the every-cycle export commit (the #2328 ~288 TOU-register-write-batches/day signature) stayed live; keep that reading for pre-#5126 logs, where counting `Successfully pressed button` lines settles the actual volume. (The every-cycle `H M` time-entity rewrite remains, deliberately — #1529 write reliability); `press_and_poll_button()` takes a `side` argument so a split-button config no longer presses the unrelated side's button (which is what cleared `timed_charge_current` alongside the intended one); and `adjust_inverter_mode()` now sleeps 30s after a real window change as a GivTCP settle workaround, so expect a 30s pause per real window change on those types. Keep the mechanism for pre-merge logs, where the signature is a button press logged every 5 minutes inside a window. Away from the CID work, `automatic_config()` had an asymmetric arg-binding defect, **fixed in PR #5308 (merged 2026-09-29, v9.3.3) - keep the pre-#5308 signature:** it bound `battery_rate_max` - which fed *both* the charge and the discharge cap - to the per-device `max_charge_power` entity, while the `max_discharge_power` entity it also creates was bound to nothing; `inverter.py`'s `min(inverter_limit_discharge, battery_rate_max_raw)` then capped discharge at the charge limit on an asymmetric inverter, and `inverter_limit_discharge` could not lift it because the raw value was the binding constraint (GH#4940, confirmed against the reporter's debug yaml). The fix: Solis publishes a `battery_rate_max` sensor holding the **larger** of the two per-direction limits and binds `battery_rate_max` to it - auto-discovery still beats a manual apps.yaml value for that binding (`component_base.py:96-110`), so an apps.yaml `battery_rate_max` does not opt out - and binds `inverter_limit_charge`/`inverter_limit_discharge` to the two per-direction limits with `overwrite=False`, so an apps.yaml `inverter_limit_*` (an AC rating, a DNO cap) still wins. Separately, on Solis Cloud/TOU-V2 the per-slot current write is clamped by `write_time_windows()` (`solis.py`, `new_current = min(slot_data['discharge_current'], cap)`). **The cap's source was rewritten in PR #5314 (merged 2026-09-30, v9.3.3) - keep the pre-#5314 signature:** the cap used to be the *minimum* `sysCommand.max` metadata across all six slot CIDs from `cached_infos` — read here at the time as deliberate, but the metadata is actually a UI definition for a *list of model codes* (its `productModel`), and the cloud hands back definitions for other models: an S5-EH1P5K-L (model 3104, 5kW, registers reading 100A) was given 60A for its slot currents from the models-3101/3102 definition and 62.5A for slot 1 charge from models 3111/3140/3145's, holding exports to ~3.6kW (GH#5068). `slot_current_limits()` now ignores the metadata: the cap starts from the battery-side limit register itself (CID 7224 charge / 7226 discharge, read as-is - not multiplied by pack count - and `SOLIS_SLOT_CURRENT_DEFAULT_AMPS` when never read), and once `ensure_slot_current_probed()` has measured what the slots really accept a **measured** ceiling replaces the rated estimate (PR #5309's probe, runnable as `python3 apps/predbat/solis.py --probe-ceiling` without `--write`; it starts from the register and the sweep ends with a half-amp step, so a 62.5A ceiling is found exactly rather than rounded down, with candidates written as the inverter shows them). The verify-echo traps survive the rewrite: the number entity is republished from the local `charge_discharge_time_windows` cache, so the write verify reads back the *clamped* value converted at the **live** battery voltage (60A × ~51.6V ≈ 3094W, not 48V) and logs `didn't complete got 3094.0` every cycle — while a "Wrote N successfully now N" line can also be self-referential (cached == clamped skips the CID write entirely). Grep `CID 5967 ... is set to` to see what actually went to the inverter. The planner models `battery_rate_max_discharge` from apps.yaml/entity max and knows nothing of the clamp, so the plan is systematically optimistic while it binds; on a current build set `inverter_limit_discharge` to the measured cap × pack voltage, pre-#5314 it was 60A × pack voltage. #4220 is the same 60A clamp from the user's angle. Separately it built every arg list - PV, load and grid included - from the battery-filtered `devices` list (`solis.py:1456`, narrowed by `cd6e7c79`), dropping a PV-only inverter's generation from `pv_today`/`pv_power`; the PV half was fixed in PR #4923, load/grid deliberately left battery-only because those registers can overlap on a shared-CT install. **A string inverter that declares its no-battery state in *neither* place was still enrolled as a battery inverter (GH#5279, fixed in PR #5281, merged 2026-09-28 — keep the pre-#5281 signature):** `_reports_no_battery()` was deliberately narrow (absence of the fields = "unknown", never dropped), and the captured string model (product `0106`) reports `batteryType '0'` — the code batteryList uses for "No Battery" on the alternative firmware, not a name — with an empty `batteryList` and zero battery readings, while `batteryHealthSoh: 0` parses as `0.0` rather than `None`, so `automatic_config()` counted it as a battery and bound control args to it (the #4707 refused-write stream). Signature: `Configuring Predbat for N inverter(s) with batteries` with **no** `Skipping inverter … reports no battery` line and **no** `Including N inverter(s) with no battery in the PV totals` line, plus `Set arg soc_percent = [two entries]`. #5281's detection requires all of: no battery type named (absent, empty or the code `0`), no batteryList entry, and `batteryVoltage` **and** `batteryCapacitySoc` both present and 0 — an unread detail or a named pack reporting zeros still counts as a battery, and a named pack is decided before the readings are looked at; a PV-only inverter now gets `inverterDetail` only (startup TOU read and hourly register reads skipped, publish stops at the detail sensors, event handlers ignore it) and takes its place in the PV totals. The #4923 PV-args split can now apply to such an inverter, which is exactly why it could not before: it never left the battery list. Pre-#5281 workaround remains `solis_inverter_sn` pinned to the battery inverter — which silently removes all PV generation from the plan, so on a current version the fix is the upgrade. One more shape worth remembering: the entity-event queue used to drain at the top of `run()`, before the `first` block created the `ClientSession` and discovered inverters, so an event queued during startup executed against `session=None` with an empty `inverter_sn` - and was popped before execution, making it a silently lost write (fixed in `9fe1f7e0`). Three September fixes and one still-live trap. **PR #5089 (merged 2026-09-14)** closed the SolisCloud API-allowance family (GH#5087/#5091): `B0115` ("Datalogger offline or disconnected") and `R0000` ("Daily API request allowance exhausted") are now classified in `SOLIS_API_CODES` and excluded from retry (`SOLIS_API_CODES_NO_RETRY`) - retrying cannot change the answer and every attempt still spends one of the 200 daily requests - with a per-inverter datalogger-offline cooldown (`datalogger_offline()`) so one offline datalogger backs off alone rather than the fleet; and `automatic_config()` is no longer first-cycle-only: `run()` retries it every cycle until it returns True (`automatic_config_done`), so an API outage at startup no longer leaves `load_today` unset until a Predbat restart (the ValueError-at-fetch symptom this produced has its own row in the symptom table). Keep the mechanism for pre-#5089 logs: B0115 was treated as rate limiting (10s sleep + retry, the `ad26f95a` "Quick rate limit hack") and any R-series code fell through to `Unknown code` with the full retry ladder - an offline datalogger drained the 200/day allowance and R0000 arrived hours later. `automatic_config()` makes no API calls of its own - it reads only `self.inverter_details` - and returns False when details are still missing, which is the retryable case. **PR #5095 (merged 2026-09-15)** fixed the amp↔watt conversion voltage (GH#5090): `get_nominal_voltage()` now prefers `solis_nominal_voltage` from apps.yaml (the only stated physical property), then a live reading above `SOLIS_HV_BATTERY_VOLTAGE` (HV packs keep the live voltage as they have since #4493), then BMS charge voltage classification (16S at/above 55V charge voltage, 15S below - settings rather than measurements, so the result holds across polls), then the 48V fallback; the live voltage had swung 10-14% over a month on every sampled system, dragging `battery_rate_max`, the write tolerance computed from it and the read-back of every rate setpoint along. No workaround existed pre-#5095 - `solis_nominal_voltage` fed only the capacity path (`get_capacity_voltage()`), so telling a user to set it fixed capacity but not the rate paths. Diagnostic tell for pre-#5095 logs: failed-write readbacks at a constant ratio of the target (e.g. exactly 66/70) are consistent with write-time vs read-time voltage difference alone. **Still live (GH#5093):** `SolisAPI.run()` folds per-inverter failures into one `poll_success` bool - discovery failure, any inverter's detail fetch, TOU window reads, the in-window re-read - and `ComponentBase.start()` gates startup on that return: False → `api_started` never set, and the startup burst (discovery, details, `startup_reset_registers()`, the infrequent poll) re-runs on a doubling backoff from 60s to a 128-minute cap, so a multi-inverter fleet is held hostage to the least reachable device and every retry re-spends quota. Per-inverter success booleans already exist at every `poll_success` site, so the fix is contained to `run()`; the interim workaround is listing only working SNs in `solis_inverter_sn`. PR #5185 (merged 2026-09-23) also backs off `startup_reset_registers()` per-inverter with a persistent B0600 pause, and a refused startup register read no longer aborts startup — the try/except now sitting at the `run()` call site covers both triggers, the #5177 B0600 refusal and the SolisCloud timeout (GH#5202, closed as its dup: the traceback matched main line-for-line, the same `SolisAPIError` escaping the same call). Three Modbus-vs-component bindings and a mode-domain split (GH#5127/#5128/#5129, triaged 2026-09-17). On `charge_time_format: "H M"` types (`GS_fb00`, most cloud subtypes) `adjust_charge_window()` rewrites the charge start/end time entities on every 5-minute cycle of a programmed window (the format check in its write condition, inverter.py) - deliberately, the same #1529 write-reliability workaround the export side has - so "Predbat hammers my Solis with writes" is this first, not SoC-hold micromanagement: `adjust_battery_target()` is change-gated (logs `already at target` and writes nothing when equal) and the hold path's `disable_charge_window()` writes nothing unless the enable switch was on. The Modbus-path types (`GS`, `GS_fb00`, `SX4`, Sofar) have **no Predbat component at all** - every control arg is bound only by the hand-edited template; `create_missing_arg()` only ever *dummies* absent capabilities (a dummy `charge_limit` 100 when `has_target_soc` is False), so no INVERTER_DEF default binds a real entity, and a GS_fb00 user who misses the template's `charge_limit` uncomment gets a silent 100.0 target plus `No entity_id for charge_limit to write N` on every target change while the inverter keeps charging to its own default (GH#5129 - the signature is the repeated `No entity_id` line, not a startup error). On the Solis Cloud component the opposite applies: `automatic_config()` binds `charge_limit` with `set_arg_auto(..., overwrite=True)` (solis.py), so an apps.yaml `charge_limit:` pointing at the inverter's own timed-charge-SoC entity is replaced and logged `auto-discovery wins for this setting` - the first thing to grep a reporter's log for when #2328's guidance "didn't help" (GH#5127). Separately, Solis has **two mode-switch control paths whose mode domains do not overlap** (GH#5128): the HA `solax_modbus` integration (`inverter_type: "GS"`) writes a mode *name* from `SOLAX_SOLIS_MODES`/`SOLAX_SOLIS_MODES_NEW` through `alt_charge_discharge_enable()` and only ever computes 33/35 - the 64/96/98 "Feed-in priority" entries are dead in that table - while the SolisCloud API path above does use feed-in priority in its idle paths; so a "why doesn't Predbat use feed-in priority" report is GH#5128's gap on modbus and already-covered on the API component. Trap for any fix there: 33/35 mean different things in the two tables (`SOLAX_SOLIS_MODES` vs `_NEW`, the `solax_modbus_new` switch), so writes must go through the table-derived name, never a bare number; the reusable pattern is `support_feedin_first` (PR #5038, Fox). Check `fox.py`/`gecloud.py` for the same fold-everything-into-one-bool shape - GE Cloud still runs its own auto-config under `if first:` only, with no retry, so its self-heal after an outage is the pre-#5089-era slow one. **GH#5187, fixed in PR #5189 (merged 2026-09-24) - keep the mechanism for pre-#5189 logs.** Slot currents are now capped at the inverter's rated power (from the `inverter_size` sensor, logged `Capping slot currents on ... at the most the inverter can deliver at its rated power`); pre-#5189 the bound was only CID 7226 × count and the per-slot `sysCommand.max` metadata, so a 3.6 kW AC inverter reading 100 A at CID 7226 refused every 100 A slot-current write and left the slot enabled at 0 A, while the planner side never saw the AC rating - pre-#5308 auto-config bound `battery_rate_max` to `max_charge_power` (GH#4940) and set no `inverter_limit_discharge`; since PR #5308 the component binds `inverter_limit_discharge` to `max_discharge_power` and an apps.yaml value still wins (`overwrite=False`), so `inverter_limit_discharge: 3600` remains the plan-side workaround where the modelling matters. Implausible recovery SoC (CID 7229) no longer defeats the #4706 clamp: 0/1/≤over-discharge/>100 now falls back to `over_discharge_soc + 1` with a log line; pre-#5189 a 65521 read passed the guard and made Predbat attempt to write the recovery register itself down every cycle. **Still live:** the write-skip on a never-read slot current CID — `cached_values.get(cid, cap)` defaults to the cap, so a register the poll has never returned reads as "no change" and **no write is attempted at all** (`solis.py`, both charge and discharge paths); if a fixture shows no write of the slot CID, check whether the fixture seeded it in `cached_values` before assuming a bug. Test notes from #5177: `test_solis.py`'s first-cycle tests mock `startup_reset_registers` via instance attribute (`del api.startup_reset_registers` restores the real method); a canned `api.session` set before `run(0, True)` reaches the real request path; and do not patch `asyncio.sleep` to speed up a retried error — `_with_retry` budgets with `time.monotonic()`, so a no-op sleep just burns the real ~30 s budget across ~8 attempts. The hold lever on Solis modbus is window-scoped (GH#5152, code-verified): the only discharge-suppression write is `set_current_from_power()`'s `timed_{charge,discharge}_current` (re-asserted on every rate adjustment, #4415); the standing `battery_discharge_current` registers are never driven anywhere in the repo, and GS makes the timed write the *only* lever (`has_reserve_soc: False`, `has_timed_pause: False`) — so a hold's 0A binds only inside the timed slot, and a "hold/freeze doesn't stop discharge" report on Solis-modbus points at GH#5152 before any plan logic: the execute-side hold gates fire correctly, the lever they pull is what is window-scoped. (The "only binds during the slot" firmware semantics is read off the register names in the solax_modbus integration, not yet live-log confirmed.) Same firmware on the **solax_modbus** path: `inverter_type: GS` with the integration's plain "Solis" plugin on V2/FB00 firmware shows as `Setting Solis Energy Control Switch to 35 Self-Use from 33 Self-Use - No Timed Charge/Discharge` every cycle, the write "succeeds", and the next read is 33 again — a `Control interference: N change(s) in the last 24h, sustained on [...energy_storage_control_switch]` count in the hundreds. Charging still worked (SoC 28→100% overnight) but exports never ran: the discharge slot's times were written but its per-slot enable (43707) stayed 0, which the `solis.py` V2 time-window decode showed as `discharge_enable: 0`. Not a code bug — the user needs the solax_modbus "Solis FB00" plugin (per-slot enable switches, no mode 35) and `inverter_type: GS_fb00` via `templates/ginlong_solis_fb00.yaml`; confirmed fixed on a live system on 26 Sep 2026. Adding Solis Cloud credentials *without* `solis_automatic: True` on such a system leaves inverter 0 bound as GS while `solis.py` still runs its slot/mode logic with no plan — it logs `no active slot`, clears discharge slot 1 and sets 'Self-Use - No Timed Charge/Discharge', fighting the modbus path. **GH#5190 (verified on main): PEP 515 underscores defeat every numeric-parse guard.** Python ≥3.6 `int()`/`float()` accept `_` digit separators, so a junk register string such as `"1234567890_00"` (datalogger off the remote-control platform; polls fail with `B0063 "Lack of iot platform three elements"`) parses as `123456789000` with no exception — `parse_cid_int()`, `cid_value_matches()` and the raw `int(...)`/`float(...)` coercion all accept it, so every `try/except ValueError` fallback silently passes. Consequences seen in the reporter's config: `reserve`/`battery_min_soc` auto-bound to the junk percent (inverter.py raises `set_reserve_min` to it), CID 636 never converging (every-minute write + failed-verify repeat), and no `B0063` handling anywhere in `solis.py`. The same gap is in the core `get_arg()` float coercion (`userinterface.py`), so it generalises past Solis: start there on any "value reads as an absurdly huge number / the fallback never fires" report. Fix shape per the issue: one strict-parse helper (plain decimal only), range checks, publish `None` for unparseable reads, never write CID 636 from an unreadable current mode. Holds and freeze (PR #5265): `reserve` is now bound to the **Battery Reserve SOC** (CID 157, `reserve_soc`) and the row declares `has_reserve_soc` True - `battery_min_soc` stays the over-discharge SOC (CID 158). Solis only treats the Battery Reserve SOC as a discharge floor while CID 636 bit 4 (`SOLIS_BIT_BACKUP_MODE`) is on, so `set_storage_mode_if_needed()` always sets that bit: `set_reserve_enable` only decides whether Predbat *changes* the reserve, and with it off `inverter.py` plans against the live reserve, so the inverter must enforce it too (gating the bit on `set_reserve_enable` left the plan flooring at a register the inverter ignored). Reserve writes are shown in the cache and entity at once by `show_soc_limit_now()` and written from the event queue - non-slot events wait for the next `run()`, far beyond Predbat's 2s write verify. Hold detection has a 1% hysteresis (`SOLIS_HOLD_HYSTERESIS`) so the mode does not flip across the target; a hold outside a slot needs no extra grid-charging logic because outside slots the component already writes a No-Grid-Charging mode. Inside an open charge slot with the battery at or above the slot's SOC (`is_holding_in_charge_slot()` - a freeze charge, or a charge that reached its target) it writes `Self-Use - No Grid Charging` (`ENUM_SELF_USE_HOLD`: 3 on V1 where bit 1 is the slot enable, 1 on V2) ahead of the 0A -> Feed-in rule, which on V1 used to select Feed-in priority - exporting PV - during a freeze charge. **`Warn: Inverter N battery_scaling read as 0.0 ... retaining last known value` repeating every cycle (GH#5368, code-verified on main + `triage_test.sh inverter`):** `solis_automatic` binds `battery_scaling` to the per-inverter `_battery_soh` sensors (`solis.py`), which publish Solis Cloud `batteryHealthSoh`/100 **as-is, including 0** (deliberate, the #4494 decision); the `refresh_config()` guard (`inverter.py`, from PR #4500) retains the last valid value (default 1.0), so **the plan is unaffected** and the report is about the repeat — a warning whose guard comment assumed a transient outage but never sees a valid read again, an enhancement (one-time/backoff per inverter), not a planning bug. apps.yaml `battery_scaling` is **not** a workaround: `set_arg_auto(..., overwrite=True)` (`component_base.py`) replaces an explicit apps.yaml value with the discovered sensor binding, so the only config-level opt-out is `solis_automatic: false`. Keep the SOH-derived `battery_scaling` distinct from `battery_scaling_auto` (usage-data degradation measurement, `fetch.py`) — different knobs, easily confused in reports. | `solis` | +| Fox (`fox.py`) | The cloud API returns errno 42015/44096 for settings a given device does not support (`FOX_SETTINGS_UNSUPPORTED_ERRNO`); those are marked unavailable and never polled or written again. Entity type matters — WorkMode is a select, ExportLimit a number. Two later capacity/schedule traps, both still live: `publish_data()` sums every `batteryList` entry unconditionally (`fox.py:1779-1781`), and an AIO ESS returns one physical pack as four `bmu` entries all carrying the inverter's own `batterySN`, so `soc_max` comes out at 4x the correct `batteryDesignCapacity` and `soc_kw` with it (GH#4919, read out of the reporter's own API response). `fox_automatic: true` re-`set_arg`s `soc_max` every cycle and `soc_max` is not in `CONFIG_API_OVERRIDE`, so apps.yaml cannot override it - the escape hatch is `fox_automatic: false`. With the Mode Scheduler deleted the scheduler `enable` flag reads 0 and `compute_schedule()` (`fox.py:1080`) derives the displayed charge window from the legacy `forceChargeTime` read, which Predbat cannot clear: `set_battery_charging_time()` (`fox.py:1009`), the only writer of that endpoint, has zero production callers - re-verified on main, where it is referenced only by `test_fox_api.py` (GH#4939). Note this is the *residual* half of that issue: the blocker the reporter actually hit was HA event routing filtering on the literal string `predbat` rather than the configured entity prefix, **fixed in PR #4962**. Unlike the self-referential HA verify trap below, a Fox `didn't complete got X` is a genuine failure - the read-back polls entities the component republishes from Fox Cloud. A separate trap on the **HA/modbus path** (inverter type `FoxESS`, not `FoxCloud`): the stock template binds `reserve:` to `number.foxess_min_soc_on_grid` (`templates/fox.yaml`), and that file's own mode table ties both Freeze charging and Hold charging to that entity. A reporter had `reserve:` bound to `number.foxess_inverter_min_soc` instead - a name that appears nowhere in `templates/` or `docs/` - so Predbat never read or wrote `min_soc_on_grid`; when the inverter stranded it at 100 (observed sitting there six days) Predbat could neither see it nor clear it, and `min_soc_on_grid` at 100% stops the battery discharging while grid-connected. `FoxESS` is `has_idle_time: False` (`config.py`), so the `idle time is ...` line is bookkeeping only and never programs a demand period - discharge outside a charge window depends entirely on the inverter's own settings, which is what makes a wrong reserve binding silent (GH#4961). FoxCloud freeze export (GH#5015, **fixed in PR #5038**, v9.0.2 - keep the mechanism for pre-v9.0.2 logs): `support_feedin_first: True` makes `prediction.py`'s freeze branch model freeze export as genuine Feed-in-First — load exports up to the limit, only the surplus beyond it charges the battery — but the cloud path used to never select the `Feedin` work mode: `apply_battery_schedule()` (`fox.py`) only emitted `SelfUse`/`ForceCharge`/`ForceDischarge` groups and `adjust_inverter_mode()` hardcodes `SelfUse` for Fox (`inverter.py`), so the inverter sat in SelfUse where surplus PV charges the battery *before* exporting and the plan's freeze-export revenue didn't match the hardware. #5038 makes `apply_battery_schedule()` select `Feedin` as the baseline work mode whenever freeze export is requested; note the cloud API's verify read is slow there — a write can log success, read back pre-write 3s later and read back the new mode after ~18s (recorded in a `fox.py` comment), so a quick verify read is not evidence of failure. `support_feedin_first` therefore means two different things depending on connection method: the modelling is correct on the FoxESS HA/modbus path, where `discharge_freeze_service` genuinely selects Feed-in First (#4207, fixed by #4425), while #4582 extended the flag to the cloud defs with no equivalent execution path. Log check: a freeze-export day shows hundreds of full-day `SelfUse` groups in `Fox: New schedule` lines and never a `ForceDischarge` or `Feedin` group — and `ForceDischarge` appearing in a `Fetch scheduler V1 returned` enum list is **not** evidence a discharge window was written. Workaround: turn `set_export_freeze` off so the planner stops selecting freeze-export slots. #4182 is the neighbouring open issue. **The device MaxSoc setting is published but never bound (GH#5395, code-read on main 2026-10-04):** `number._fox__setting_maxsoc` is user-writable with a working write-back to Fox Cloud, but `automatic_config()` binds only `soc_max`/`reserve`/`battery_min_soc`/`export_limit`, there is no per-inverter max-SoC arg in `inverter.py`, and the planner's only upper ceiling is the global `best_soc_max` in kWh — so "the plan models full capacity while the inverter stops at the device's Max SoC" is a real gap; a tree-wide grep for `maxsoc` landing only in the component (plus its test) is the tell. The writer hardcodes `maxSoc: 100` on baseline/filler groups and sends `max(soc, reserve)` on charge groups, while schedule reads derive the published entity with **max across groups** — the reporter's device read back 80 against writes of 100, implying the device applies Max SoC on that write path somewhere (**unverified**). Family: #4832 (GE Cloud floor-side sibling), #2128 (the same ask as a user input number), #3837 (FoxESS modbus: the plan charge target is never written to the device max_soc entity), #3398 (`fdSoc: 100`); do not confuse with `manual_soc_max` (a time-of-day manual percent override, PR #4841). | `fox_api`, `fox_oauth` | +| Solis (`solis.py`) | `SOLIS_CID_STORAGE_MODE = 636` is Modbus 43110, a bit mask. On firmware "4B and above" (what `is_tou_v2_mode()` detects: CID 6798 reads 43605) the timed charge/discharge enable moved to the per-slot registers — `SOLIS_CID_CHARGE_ENABLE_BASE`/`..._DISCHARGE_ENABLE_BASE`, Modbus 43707, six slots not three — and every mode value carrying TOU bit 1 (3/35/43/51/98) was dropped; 35 became 33, 98 became 96. Such an inverter answers a CID 636 write with code 0 and reads back without bit 1, so a log full of CID 636 verification warnings is a refused bit, not a failed write (GH#4707). `set_storage_mode_if_needed()` decides the bit from `is_tou_v2_mode()`: never asked for on V2, always asked for and retried on V1, where bit 1 still *is* the timed charge/discharge enable. Do not try to learn it from the read-back instead — GH#4710 did, and GH#4774 showed why that cannot work: on six inverters the same write of 179 verifies minutes before and minutes after the one that reads back 177, so a single post-write read is not evidence of a firmware property. Latching it was also silent rather than noisy: the verify read refreshes `cached_values`, so after a stripped write the cache held the stripped value, the suppressed computed value matched it, and `set_storage_mode_if_needed()` stopped writing CID 636 at all for the whole 8-hour verdict — across an overnight charge window. GH#4239's "only retained while a window is configured" is not the whole story — it was refused with slot 1 enabled and its window in force. `read_and_write_cid()` re-reads once after `verify_settle_seconds` before calling a mismatch a failure, because the immediate verify read is taken about half a second after the write. On the SolisCloud select path specifically, **PR #5247 (merged 2026-09-26) makes a discovered inverter's slot-time writes run their handler immediately** so Predbat's read-back sees them in the same cycle — pre-#5247 the select callback only queued the event (GH#4875) and the queue drained at the top of `run()` up to a minute later, so the write verified against the stale value and logged a spurious `didn't complete` once per moved slot time (pass 1, V1 and the slot SoC/current writes keep their every-minute retry). Do not "stop the failed verify poisoning the cache" — `write_cid()` caches the value it *requested* and the post-write `read_cid()` deliberately overwrites it with what the inverter actually reports, so the cache mirrors the inverter rather than Predbat's intent; skipping that update would leave the cache agreeing with the write that just failed, and change detection would then never retry it. The two control paths choose the storage mode differently and it matters: V1 picks from `in_charge_slot`/`in_discharge_slot`, which are clock tests, so CID 636 changes value and gets written the moment a window opens; V2 picks from `slot1_active`, which is only "slot 1 has a window configured", so the mode is written when the slot is *programmed* and nothing at all happens at the window boundary. On the GH#4774 night the last CID 636 write was 22 minutes before the window opened, and the comparison night that worked had one near the window - so `claim_window_mode_assertion()` now asserts the mode once per window on both paths. Slot registers are polled hourly, plus every 5 minutes while `is_inside_active_window()` is true, so drift during the window that matters is caught in the same cycle it happens. GH#4774 also left a useful negative result: that inverter grid-charged from 20% to 91% overnight with bit 1 clear throughout, confirming the per-slot enables alone drive timed charging on V2 — so a "charge window never started" report on V2 firmware is not a CID 636 problem and needs looking at elsewhere. Separately, an inverter reporting `batteryType 'No Battery'` stays in `self.inverter_sn` for polling; `is_battery_inverter()` keeps control writes off it, which was most of the warning volume in GH#4707. That explanation doesn't cover every CID 636 report, though: a later GH#4707 comment showed zero `setting storage mode to` idle-mode log lines, and both V1 mode-decision idle paths always log that line — so a report with neither line cannot be coming from `write_time_windows_if_changed()` at all. The only remaining write path is the HA select handler, `set_storage_mode_value()` at `solis.py:2574`. Check which write path actually ran before assuming this entry's explanation applies. Separately, on inverters using `H M` time format (`GS_fb00`, most cloud subtypes) `adjust_force_export()` used to re-commit a stable export window every cycle — `is_hm_format` alone made `changed_start_end` true, so `press_and_poll_button()` fired on every cycle of an unchanged export window (GH#4709, confirmed live on main at the time) — and PR #4713 fixed the twin idle-cycle press where the times being managed came back as `None` and never compared equal to what the inverter still reported (GH#4712, #2328). **Both merged together in PR #4711 (`bc853a0e`, 2026-09-12, post-v9.0.2)**: `adjust_force_export()` re-commits only on a real schedule change through the commit-once ledger (`last_committed`/`commit_pending`, `commit_needed()`/`record_commit()` in `inverter.py` — the attribute this entry used to name, `last_export_schedule_committed`, is gone), and a commit is only recorded when the writes *and* the button press both succeeded; **the guard only became real when PR #5126 (merged 2026-09-26, unreleased at the time of writing) stopped rebuilding Inverter objects every cycle** — pre-#5126 the rebuild reset the latch to empty each cycle, so the guard was decorative and the every-cycle export commit (the #2328 ~288 TOU-register-write-batches/day signature) stayed live; keep that reading for pre-#5126 logs, where counting `Successfully pressed button` lines settles the actual volume. (The every-cycle `H M` time-entity rewrite remains, deliberately — #1529 write reliability); `press_and_poll_button()` takes a `side` argument so a split-button config no longer presses the unrelated side's button (which is what cleared `timed_charge_current` alongside the intended one); and `adjust_inverter_mode()` now sleeps 30s after a real window change as a GivTCP settle workaround, so expect a 30s pause per real window change on those types. Keep the mechanism for pre-merge logs, where the signature is a button press logged every 5 minutes inside a window. Away from the CID work, `automatic_config()` had an asymmetric arg-binding defect, **fixed in PR #5308 (merged 2026-09-29, v9.3.3) - keep the pre-#5308 signature:** it bound `battery_rate_max` - which fed *both* the charge and the discharge cap - to the per-device `max_charge_power` entity, while the `max_discharge_power` entity it also creates was bound to nothing; `inverter.py`'s `min(inverter_limit_discharge, battery_rate_max_raw)` then capped discharge at the charge limit on an asymmetric inverter, and `inverter_limit_discharge` could not lift it because the raw value was the binding constraint (GH#4940, confirmed against the reporter's debug yaml). The fix: Solis publishes a `battery_rate_max` sensor holding the **larger** of the two per-direction limits and binds `battery_rate_max` to it - auto-discovery still beats a manual apps.yaml value for that binding (`component_base.py:96-110`), so an apps.yaml `battery_rate_max` does not opt out - and binds `inverter_limit_charge`/`inverter_limit_discharge` to the two per-direction limits with `overwrite=False`, so an apps.yaml `inverter_limit_*` (an AC rating, a DNO cap) still wins. Separately, on Solis Cloud/TOU-V2 the per-slot current write is clamped by `write_time_windows()` (`solis.py`, `new_current = min(slot_data['discharge_current'], cap)`). **The cap's source was rewritten in PR #5314 (merged 2026-09-30, v9.3.3) - keep the pre-#5314 signature:** the cap used to be the *minimum* `sysCommand.max` metadata across all six slot CIDs from `cached_infos` — read here at the time as deliberate, but the metadata is actually a UI definition for a *list of model codes* (its `productModel`), and the cloud hands back definitions for other models: an S5-EH1P5K-L (model 3104, 5kW, registers reading 100A) was given 60A for its slot currents from the models-3101/3102 definition and 62.5A for slot 1 charge from models 3111/3140/3145's, holding exports to ~3.6kW (GH#5068). `slot_current_limits()` now ignores the metadata: the cap starts from the battery-side limit register itself (CID 7224 charge / 7226 discharge, read as-is - not multiplied by pack count - and `SOLIS_SLOT_CURRENT_DEFAULT_AMPS` when never read), and once `ensure_slot_current_probed()` has measured what the slots really accept a **measured** ceiling replaces the rated estimate (PR #5309's probe, runnable as `python3 apps/predbat/solis.py --probe-ceiling` without `--write`; it starts from the register and the sweep ends with a half-amp step, so a 62.5A ceiling is found exactly rather than rounded down, with candidates written as the inverter shows them). The verify-echo traps survive the rewrite: the number entity is republished from the local `charge_discharge_time_windows` cache, so the write verify reads back the *clamped* value converted at the **live** battery voltage (60A × ~51.6V ≈ 3094W, not 48V) and logs `didn't complete got 3094.0` every cycle — while a "Wrote N successfully now N" line can also be self-referential (cached == clamped skips the CID write entirely). Grep `CID 5967 ... is set to` to see what actually went to the inverter. The planner models `battery_rate_max_discharge` from apps.yaml/entity max and knows nothing of the clamp, so the plan is systematically optimistic while it binds; on a current build set `inverter_limit_discharge` to the measured cap × pack voltage, pre-#5314 it was 60A × pack voltage. #4220 is the same 60A clamp from the user's angle. Separately it built every arg list - PV, load and grid included - from the battery-filtered `devices` list (`solis.py:1456`, narrowed by `cd6e7c79`), dropping a PV-only inverter's generation from `pv_today`/`pv_power`; the PV half was fixed in PR #4923, load/grid deliberately left battery-only because those registers can overlap on a shared-CT install. **A string inverter that declares its no-battery state in *neither* place was still enrolled as a battery inverter (GH#5279, fixed in PR #5281, merged 2026-09-28 — keep the pre-#5281 signature):** `_reports_no_battery()` was deliberately narrow (absence of the fields = "unknown", never dropped), and the captured string model (product `0106`) reports `batteryType '0'` — the code batteryList uses for "No Battery" on the alternative firmware, not a name — with an empty `batteryList` and zero battery readings, while `batteryHealthSoh: 0` parses as `0.0` rather than `None`, so `automatic_config()` counted it as a battery and bound control args to it (the #4707 refused-write stream). Signature: `Configuring Predbat for N inverter(s) with batteries` with **no** `Skipping inverter … reports no battery` line and **no** `Including N inverter(s) with no battery in the PV totals` line, plus `Set arg soc_percent = [two entries]`. #5281's detection requires all of: no battery type named (absent, empty or the code `0`), no batteryList entry, and `batteryVoltage` **and** `batteryCapacitySoc` both present and 0 — an unread detail or a named pack reporting zeros still counts as a battery, and a named pack is decided before the readings are looked at; a PV-only inverter now gets `inverterDetail` only (startup TOU read and hourly register reads skipped, publish stops at the detail sensors, event handlers ignore it) and takes its place in the PV totals. The #4923 PV-args split can now apply to such an inverter, which is exactly why it could not before: it never left the battery list. Pre-#5281 workaround remains `solis_inverter_sn` pinned to the battery inverter — which silently removes all PV generation from the plan, so on a current version the fix is the upgrade. One more shape worth remembering: the entity-event queue used to drain at the top of `run()`, before the `first` block created the `ClientSession` and discovered inverters, so an event queued during startup executed against `session=None` with an empty `inverter_sn` - and was popped before execution, making it a silently lost write (fixed in `9fe1f7e0`). Three September fixes and one still-live trap. **PR #5089 (merged 2026-09-14)** closed the SolisCloud API-allowance family (GH#5087/#5091): `B0115` ("Datalogger offline or disconnected") and `R0000` ("Daily API request allowance exhausted") are now classified in `SOLIS_API_CODES` and excluded from retry (`SOLIS_API_CODES_NO_RETRY`) - retrying cannot change the answer and every attempt still spends one of the 200 daily requests - with a per-inverter datalogger-offline cooldown (`datalogger_offline()`) so one offline datalogger backs off alone rather than the fleet; and `automatic_config()` is no longer first-cycle-only: `run()` retries it every cycle until it returns True (`automatic_config_done`), so an API outage at startup no longer leaves `load_today` unset until a Predbat restart (the ValueError-at-fetch symptom this produced has its own row in the symptom table). Keep the mechanism for pre-#5089 logs: B0115 was treated as rate limiting (10s sleep + retry, the `ad26f95a` "Quick rate limit hack") and any R-series code fell through to `Unknown code` with the full retry ladder - an offline datalogger drained the 200/day allowance and R0000 arrived hours later. `automatic_config()` makes no API calls of its own - it reads only `self.inverter_details` - and returns False when details are still missing, which is the retryable case. **PR #5095 (merged 2026-09-15)** fixed the amp↔watt conversion voltage (GH#5090): `get_nominal_voltage()` now prefers `solis_nominal_voltage` from apps.yaml (the only stated physical property), then a live reading above `SOLIS_HV_BATTERY_VOLTAGE` (HV packs keep the live voltage as they have since #4493), then BMS charge voltage classification (16S at/above 55V charge voltage, 15S below - settings rather than measurements, so the result holds across polls), then the 48V fallback; the live voltage had swung 10-14% over a month on every sampled system, dragging `battery_rate_max`, the write tolerance computed from it and the read-back of every rate setpoint along. No workaround existed pre-#5095 - `solis_nominal_voltage` fed only the capacity path (`get_capacity_voltage()`), so telling a user to set it fixed capacity but not the rate paths. Diagnostic tell for pre-#5095 logs: failed-write readbacks at a constant ratio of the target (e.g. exactly 66/70) are consistent with write-time vs read-time voltage difference alone. **Still live (GH#5093):** `SolisAPI.run()` folds per-inverter failures into one `poll_success` bool - discovery failure, any inverter's detail fetch, TOU window reads, the in-window re-read - and `ComponentBase.start()` gates startup on that return: False → `api_started` never set, and the startup burst (discovery, details, `startup_reset_registers()`, the infrequent poll) re-runs on a doubling backoff from 60s to a 128-minute cap, so a multi-inverter fleet is held hostage to the least reachable device and every retry re-spends quota. Per-inverter success booleans already exist at every `poll_success` site, so the fix is contained to `run()`; the interim workaround is listing only working SNs in `solis_inverter_sn`. PR #5185 (merged 2026-09-23) also backs off `startup_reset_registers()` per-inverter with a persistent B0600 pause, and a refused startup register read no longer aborts startup — the try/except now sitting at the `run()` call site covers both triggers, the #5177 B0600 refusal and the SolisCloud timeout (GH#5202, closed as its dup: the traceback matched main line-for-line, the same `SolisAPIError` escaping the same call). Three Modbus-vs-component bindings and a mode-domain split (GH#5127/#5128/#5129, triaged 2026-09-17). On `charge_time_format: "H M"` types (`GS_fb00`, most cloud subtypes) `adjust_charge_window()` rewrites the charge start/end time entities on every 5-minute cycle of a programmed window (the format check in its write condition, inverter.py) - deliberately, the same #1529 write-reliability workaround the export side has - so "Predbat hammers my Solis with writes" is this first, not SoC-hold micromanagement: `adjust_battery_target()` is change-gated (logs `already at target` and writes nothing when equal) and the hold path's `disable_charge_window()` writes nothing unless the enable switch was on. The Modbus-path types (`GS`, `GS_fb00`, `SX4`, Sofar) have **no Predbat component at all** - every control arg is bound only by the hand-edited template; `create_missing_arg()` only ever *dummies* absent capabilities (a dummy `charge_limit` 100 when `has_target_soc` is False), so no INVERTER_DEF default binds a real entity, and a GS_fb00 user who misses the template's `charge_limit` uncomment gets a silent 100.0 target plus `No entity_id for charge_limit to write N` on every target change while the inverter keeps charging to its own default (GH#5129 - the signature is the repeated `No entity_id` line, not a startup error). On the Solis Cloud component the opposite applies: `automatic_config()` binds `charge_limit` with `set_arg_auto(..., overwrite=True)` (solis.py), so an apps.yaml `charge_limit:` pointing at the inverter's own timed-charge-SoC entity is replaced and logged `auto-discovery wins for this setting` - the first thing to grep a reporter's log for when #2328's guidance "didn't help" (GH#5127). Separately, Solis has **two mode-switch control paths whose mode domains do not overlap** (GH#5128): the HA `solax_modbus` integration (`inverter_type: "GS"`) writes a mode *name* from `SOLAX_SOLIS_MODES`/`SOLAX_SOLIS_MODES_NEW` through `alt_charge_discharge_enable()` and only ever computes 33/35 - the 64/96/98 "Feed-in priority" entries are dead in that table - while the SolisCloud API path above does use feed-in priority in its idle paths; so a "why doesn't Predbat use feed-in priority" report is GH#5128's gap on modbus and already-covered on the API component. Trap for any fix there: 33/35 mean different things in the two tables (`SOLAX_SOLIS_MODES` vs `_NEW`, the `solax_modbus_new` switch), so writes must go through the table-derived name, never a bare number; the reusable pattern is `support_feedin_first` (PR #5038, Fox). Check `fox.py`/`gecloud.py` for the same fold-everything-into-one-bool shape - GE Cloud still runs its own auto-config under `if first:` only, with no retry, so its self-heal after an outage is the pre-#5089-era slow one. **GH#5187, fixed in PR #5189 (merged 2026-09-24) - keep the mechanism for pre-#5189 logs.** Slot currents are now capped at the inverter's rated power (from the `inverter_size` sensor, logged `Capping slot currents on ... at the most the inverter can deliver at its rated power`); pre-#5189 the bound was only CID 7226 × count and the per-slot `sysCommand.max` metadata, so a 3.6 kW AC inverter reading 100 A at CID 7226 refused every 100 A slot-current write and left the slot enabled at 0 A, while the planner side never saw the AC rating - pre-#5308 auto-config bound `battery_rate_max` to `max_charge_power` (GH#4940) and set no `inverter_limit_discharge`; since PR #5308 the component binds `inverter_limit_discharge` to `max_discharge_power` and an apps.yaml value still wins (`overwrite=False`), so `inverter_limit_discharge: 3600` remains the plan-side workaround where the modelling matters. Implausible recovery SoC (CID 7229) no longer defeats the #4706 clamp: 0/1/≤over-discharge/>100 now falls back to `over_discharge_soc + 1` with a log line; pre-#5189 a 65521 read passed the guard and made Predbat attempt to write the recovery register itself down every cycle. **Still live:** the write-skip on a never-read slot current CID — `cached_values.get(cid, cap)` defaults to the cap, so a register the poll has never returned reads as "no change" and **no write is attempted at all** (`solis.py`, both charge and discharge paths); if a fixture shows no write of the slot CID, check whether the fixture seeded it in `cached_values` before assuming a bug. Test notes from #5177: `test_solis.py`'s first-cycle tests mock `startup_reset_registers` via instance attribute (`del api.startup_reset_registers` restores the real method); a canned `api.session` set before `run(0, True)` reaches the real request path; and do not patch `asyncio.sleep` to speed up a retried error — `_with_retry` budgets with `time.monotonic()`, so a no-op sleep just burns the real ~30 s budget across ~8 attempts. The hold lever on Solis modbus is window-scoped (GH#5152, code-verified): the only discharge-suppression write is `set_current_from_power()`'s `timed_{charge,discharge}_current` (re-asserted on every rate adjustment, #4415); the standing `battery_discharge_current` registers are never driven anywhere in the repo, and GS makes the timed write the *only* lever (`has_reserve_soc: False`, `has_timed_pause: False`) — so a hold's 0A binds only inside the timed slot, and a "hold/freeze doesn't stop discharge" report on Solis-modbus points at GH#5152 before any plan logic: the execute-side hold gates fire correctly, the lever they pull is what is window-scoped. (The "only binds during the slot" firmware semantics is read off the register names in the solax_modbus integration, not yet live-log confirmed.) Same firmware on the **solax_modbus** path: `inverter_type: GS` with the integration's plain "Solis" plugin on V2/FB00 firmware shows as `Setting Solis Energy Control Switch to 35 Self-Use from 33 Self-Use - No Timed Charge/Discharge` every cycle, the write "succeeds", and the next read is 33 again — a `Control interference: N change(s) in the last 24h, sustained on [...energy_storage_control_switch]` count in the hundreds. Charging still worked (SoC 28→100% overnight) but exports never ran: the discharge slot's times were written but its per-slot enable (43707) stayed 0, which the `solis.py` V2 time-window decode showed as `discharge_enable: 0`. Not a code bug — the user needs the solax_modbus "Solis FB00" plugin (per-slot enable switches, no mode 35) and `inverter_type: GS_fb00` via `templates/ginlong_solis_fb00.yaml`; confirmed fixed on a live system on 26 Sep 2026. Adding Solis Cloud credentials *without* `solis_automatic: True` on such a system leaves inverter 0 bound as GS while `solis.py` still runs its slot/mode logic with no plan — it logs `no active slot`, clears discharge slot 1 and sets 'Self-Use - No Timed Charge/Discharge', fighting the modbus path. **GH#5190 (verified on main): PEP 515 underscores defeat every numeric-parse guard.** Python ≥3.6 `int()`/`float()` accept `_` digit separators, so a junk register string such as `"1234567890_00"` (datalogger off the remote-control platform; polls fail with `B0063 "Lack of iot platform three elements"`) parses as `123456789000` with no exception — `parse_cid_int()`, `cid_value_matches()` and the raw `int(...)`/`float(...)` coercion all accept it, so every `try/except ValueError` fallback silently passes. Consequences seen in the reporter's config: `reserve`/`battery_min_soc` auto-bound to the junk percent (inverter.py raises `set_reserve_min` to it), CID 636 never converging (every-minute write + failed-verify repeat), and no `B0063` handling anywhere in `solis.py`. The same gap is in the core `get_arg()` float coercion (`userinterface.py`), so it generalises past Solis: start there on any "value reads as an absurdly huge number / the fallback never fires" report. Fix shape per the issue: one strict-parse helper (plain decimal only), range checks, publish `None` for unparseable reads, never write CID 636 from an unreadable current mode. Holds and freeze (PR #5265): `reserve` is now bound to the **Battery Reserve SOC** (CID 157, `reserve_soc`) and the row declares `has_reserve_soc` True - `battery_min_soc` stays the over-discharge SOC (CID 158). Solis only treats the Battery Reserve SOC as a discharge floor while CID 636 bit 4 (`SOLIS_BIT_BACKUP_MODE`) is on, so `set_storage_mode_if_needed()` always sets that bit: `set_reserve_enable` only decides whether Predbat *changes* the reserve, and with it off `inverter.py` plans against the live reserve, so the inverter must enforce it too (gating the bit on `set_reserve_enable` left the plan flooring at a register the inverter ignored). Reserve writes are shown in the cache and entity at once by `show_soc_limit_now()` and written from the event queue - non-slot events wait for the next `run()`, far beyond Predbat's 2s write verify. Hold detection has a 1% hysteresis (`SOLIS_HOLD_HYSTERESIS`) so the mode does not flip across the target; a hold outside a slot needs no extra grid-charging logic because outside slots the component already writes a No-Grid-Charging mode. Inside an open charge slot with the battery at or above the slot's SOC (`is_holding_in_charge_slot()` - a freeze charge, or a charge that reached its target) it writes `Self-Use - No Grid Charging` (`ENUM_SELF_USE_HOLD`: 3 on V1 where bit 1 is the slot enable, 1 on V2) ahead of the 0A -> Feed-in rule, which on V1 used to select Feed-in priority - exporting PV - during a freeze charge. **`Warn: Inverter N battery_scaling read as 0.0 ... retaining last known value` repeating every cycle (GH#5368, code-verified on main + `triage_test.sh inverter`):** `solis_automatic` binds `battery_scaling` to the per-inverter `_battery_soh` sensors (`solis.py`), which publish Solis Cloud `batteryHealthSoh`/100 **as-is, including 0** (deliberate, the #4494 decision); the `refresh_config()` guard (`inverter.py`, from PR #4500) retains the last valid value (default 1.0), so **the plan is unaffected** and the report is about the repeat — a warning whose guard comment assumed a transient outage but never sees a valid read again, an enhancement (one-time/backoff per inverter), not a planning bug. apps.yaml `battery_scaling` is **not** a workaround: `set_arg_auto(..., overwrite=True)` (`component_base.py`) replaces an explicit apps.yaml value with the discovered sensor binding, so the only config-level opt-out is `solis_automatic: false`. Keep the SOH-derived `battery_scaling` distinct from `battery_scaling_auto` (usage-data degradation measurement, `fetch.py`) — different knobs, easily confused in reports. **CID 499 (max export power) is scaled asymmetrically (GH#5385, code-read + debug replay on main, 2026-10-04):** the read treats raw < 200 as 100 W units and multiplies by 100, publishing raw ≥ 200 as watts (`_discovery_export_limit()` documents and duplicates the split into the discovery record), while the write path (`number_event_handler()`) is unconditional `int(value) // 100` — so a register reported ≥ 200 that is actually still in 100 W units (reporter: raw 480 ⇒ a 0.48 kW export limit on both inverters) lies by 100x on every read-back; the ≥200⇒watts half came in with commit `05514917` (v8.37.4, no linked issue) and no recorded case of a Solis reporting raw watts exists, so an unconditional ×100 on the read (both call sites plus the test cases) is the maintainer-facing fix direction — but `test_solis.py::test_publish_entities_export_power_unit_conversion` pins the threshold split (case 6 asserts raw 200 stays 200.0), so a suite green there is not unit-correctness evidence, and any future "Solis export limit wrong by 100x" report with raw ≥ 200 (a ≥ 20000 W limit) maps here (related #3700). The issue's "cannot be overridden in apps.yaml" half is by design: `export_limit` binds via `set_arg_auto(..., overwrite=True)` (grep `apps_yaml_override_warned` in the dump), with `inverter_limit_charge`/`_discharge` the `overwrite=False` exceptions. | `solis` | | Solis Modbus (`GS`, `GS_fb00` in `inverter.py`) | The Energy Storage Control Switch (Modbus 43110, `energy_control_switch`) is driven by `alt_charge_discharge_enable()`, gated on the `has_solis_energy_control` capability. `GS` reaches it through `mimic_target_soc()` and uses the Timed Charge/Discharge bit as the charge enable (35/33). `GS_fb00` has a target SoC, so it never reached `mimic_target_soc()` - until the fix it never wrote the switch at all, and a switch left on `Self-Use - No Grid Charging` (1, bit 5 clear) made every charge slot hold the battery at 0 W while Predbat logged a correct slot, target and rate. For FB00 it is driven through `write_solis_energy_control()` from `adjust_charge_immediate()` on non-exporting cycles - `Backup/Reserve` (49) to charge or idle, `Backup/Reserve - No Grid Charging` (17) for a freeze or a car/iBoost/reserve hold - and from `adjust_export_immediate()` on exporting cycles - `Feed-in priority - No Grid Charging` (64) for a freeze export, `Backup/Reserve` for a real export (PR #5258). The Battery Reserve bit (bit 4) is what makes the template's `reserve` (`backup_mode_soc`, 43024) a discharge floor, so the reserve holds in `execute_plan()` work unchanged; grid charging goes off during holds because the reserve is raised to SoC + 1 and the inverter would import to reach it, again each cycle. `execute_plan()` never calls both immediate methods with a real target in one cycle, and the idle `adjust_export_immediate(100)` deliberately leaves the switch alone; GS still writes from `mimic_target_soc()`. Adding a writer anywhere else (e.g. `adjust_battery_target()`) flips the switch twice per cycle. Older logs: PR #5256 (merged 2026-09-26) held it on `Self-Use` (33) from `adjust_battery_target()`, and before that GS_fb00 never wrote the switch at all. A SolisCloud session is the usual way the switch ends up at 1: `solis.py` maps `"Self-Use - No Timed Charge/Discharge"` to `ENUM_SELF_USE_NO_GRID_CHARGING`, which clears bit 5, and writes it whenever no slot is active - so stopping the cloud component outside a slot leaves grid charging disallowed. The FB00 Solax Modbus plugin's option names differ from the pre-FB00 plugin's (`Self-Use` is 33 there, 35 here) - `SOLAX_SOLIS_MODES_FB00`. Verified from a user log (26 Sep 2026) and the plugin source. | `./run_all --test inverter` (`solis_fb00_*`, `solis_gs_*`) | | SolaX (`solax.py`) | Code `10402` is a token/auth failure, retried in-request rather than waiting for the next cycle. SolaX clamps battery minimum SOC at 10% (`SOLAX_MIN_RESERVE_PERCENT`), so `battery_min_soc` is auto-configured to stop Predbat writing limits the inverter will reject. **Follow-up on the GH#4993 thread (2026-09-21, from a 23-hour reporter log):** `10402` repeating every cycle with data still flowing is **server-side token invalidation between cycles**, not broken auth or lost data — Predbat's in-request refresh recovers every cycle. The tell: the same token value succeeds mid-cycle and is rejected next cycle while Predbat holds it untouched in memory between cycles (the auth-error branch of `_request_get_impl()` clears it, and `request_wrapper()` retries with a fresh token instead of failing the fetch), and a stated ~30-day `expires_in` against a real ~1-cycle token lifetime. Prime suspect is a second client sharing the same `solax_client_id`/`solax_client_secret` — the reporter ran the separate SolaX Developer API HA integration — **suspected, not proven** (one-active-token-per-app is common but could not be confirmed against SolaX docs); the decisive test is disabling the second client for an hour or registering each client its own app. Count `error 10402` lines against `Fetched ... successfully` lines before calling data lost, and note each 10402 also increments the component's lifetime error counter (`solax.py` auth-error branch), which is what accumulates in the Components panel — do not "fix" this by removing the in-request retry. Secondary, same thread: `Warn: SolarAPI: No solar forecast data was returned from HA sensors.` with the SolaX Cloud template means **no PV forecast source is configured at all** — with `forecast_solar`/`open_meteo`/`solcast_host`+`solcast_api_key` all absent, `fetch_pv_forecast()` falls through to the HA-sensor branch (`solcast.py` sets `active_source = "ha_sensors"`) and the template's `pv_forecast_today`-style `re:` regexes (`templates/solax_cloud.yaml`) match nothing unless the Solcast HA integration is installed. Check which source the log names ("Obtaining solar forecast from Forecast Solar API" vs "Using Solcast integration from inside HA") — a working source dropped during a template migration looks exactly like this. **Plant `totalYield` overstates PV on a mixed plant (GH#5356, code-verified on main 2026-10-02, payload semantics rest on the reporter's plant+device captures):** `automatic_config()` binds `pv_today` to the plant-level `sensor._solax__total_yield` (solax.py:447), `publish_plant_info()` fills that sensor straight from plant realtime `totalYield` (solax.py:2640), and `total_load` adds the same `totalYield` to imported+discharged−exported−charged (solax.py:2709) — but SolaX's plant total is PV **plus** every deviceType-1 inverter's AC output, so on a plant mixing a PV-only inverter with a separate AC-coupled battery inverter, battery discharge is counted twice (once in `totalDischarged`, once inside `totalYield`), PV history is inflated (reporter's calibration scaled 2.89x, ~1.6x legitimate), and the class presents as "calibration too high / overnight under-charging / wrong savings". `pv_power` is unaffected (per-device `pvMap`/`mpptMap` sums, solax.py:2276-2285). The ingredient a fix needs all already exist: a device-level `...__total_yield` sensor (solax.py:2366), failed reads keep the previous value (solax.py:2245), and `plant_inverters[plant_id]` holds every deviceType-1 device (solax.py:1483). Not a regression — the plant wiring dates to original SolaX Cloud development (#3140). The reporter's proposed fix direction (maintainer call pending): sum device `totalYield` over `plant_inverters` into a new `pv_today` sensor; caveat for whoever builds it — the plant **lifetime** total does not decompose as PV+discharge on that site, so derive from the daily relationship (which matched exactly on three days), not the lifetime number. **PR #5357:** `pv_today` is now `..._pv_yield`, summed by `get_plant_pv_yield()` from each inverter's **last known** `totalYield` (kept per inverter, seeded after a restart from its own `...__total_yield` sensor), so a failed read or a dead inverter does not drop out of the sum; an inverter with no `totalYield` of its own counts as 0, and the value is held only when an inverter that `onlineStatus` says should answer failed its read and has never given a yield. `pv_power`/`grid_power` are plant sensors summed over every inverter, no longer the first inverter listed. Device `dailyYield` reads 0.0 once the inverter sleeps in the evening, so it is no use as a daily counter. On a live plant of two DC-coupled hybrids with panels (2026-10-04) device `dailyYield` (23.3, 2.6) and `dailyACOutput` (24.1, 9.5) were separate counters and the plant `dailyYield` (33.6) was exactly the sum of the AC outputs, so the plant figure is inverter AC output on every plant shape seen and device yield is not; device yield was not checked against an independent PV measurement. Still wrong after the fix: a plant whose only inverter is an X1-AC (no `totalYield`, empty `mpptMap`/`pvMap`) falls back to the plant figure, which there is battery discharge (`dailyYield` = `dailyACOutput` = `dailyDischarged`), and its `load_power` goes negative when a third-party PV inverter SolaX cannot see is exporting. `plant_inverters` is keyed by the plant ID as SolaX gives it while entity names use the lowercased form - look up with the raw ID. **GH#5388 (2026-10-04):** a plant lifetime counter can dip and recover within a day (`totalYield` 4437.8 → 4260.4 → 4441.7, `totalDischarged` with it) and, published as received, the recovery was recorded as 181.3 kWh of PV that day - an absurd `pv_today` or `export_today` on one day followed by a PV forecast several times the array size is this, via calibration. An upward step of thousands of kWh on the same counters did not show in `pv_today`. `hold_counter_dip()` now holds the five plant totals and each inverter's `total_yield` at the last published value and takes a lower value as a genuine reset only after `SOLAX_COUNTER_DIP_HOLD_HOURS`; how long a real dip lasts was not measured. After a restart the last value is seeded from the sensor, or from the `counters` entry in storage when the sensor is gone (Predbat's sensors do not survive a Home Assistant restart); the time a dip started is saved there too, without it a restart began the hold again and a system restarted daily would never accept a genuine reset. A plant where no inverter has PV inputs (X1-AC only: no `totalYield`, empty `mpptMap`/`pvMap`) now publishes **0** for `pv_yield` instead of the plant total, with a one-off `has no inverter with PV inputs` warning; `total_load` then has no PV term, so on a site with third-party PV it under-reads by what that PV supplies and can fall while the site exports. Nothing in the SolaX data measures that PV: a 12 hour dump of `/openapi/v2/device/history_data` (params `snList`, `deviceType`, `startTime`/`endTime` in ms, `timeInterval` minutes, `businessType`) carried the same fields as realtime, yield null and `gridPowerM2` 0.0 throughout daylight, and no plant checked had a `deviceType=3` meter device. `M2` fields are meter 2, positive = export per the SolaX API text; `solax.py` does not read them, and a meter 2 that was once fitted shows as a non-zero `totalExportEnergyM2` that no longer moves. A hybrid with no panels still reports a populated `mpptMap` of zeros, so it counts as having PV inputs and is not covered by the 0. | `solax` | -| Sigenergy (`sigenergy.py`) | Lifetime/history totals are cumulative server-side and reset around EU midnight, which showed up as an overnight dip; `fetch_history_totals` applies a monotonic clamp. "Energy totals went backwards overnight" starts here. Separately, a user `apps.yaml` override of `inverter.has_reserve_soc: true` (stock `SIG` default is `False`, deliberately — reverted in `5bfc5c80` after #2873/#3124) bypasses two guards that normally make reserve writes inert for Sigenergy (`execute.py:947-950`, `inverter.py:547-549`), so `adjust_reserve()` writes straight to `number.sigen_plant_ess_discharge_cut_off_state_of_charge`. Confirmed from a reporter's log: zero occurrences of the "Inverter does not support reserve" disable message the first guard would normally log, plus 56 reserve writes landing on that register (GH#4728). The stock template's `reserve:` line also points at that same entity rather than `ess_backup_state_of_charge`, which is what makes the override immediately destructive rather than merely inert. `Applying mode=eco ... duration=720min` is emitted only by the no-window fall-through branch (`sigenergy.py:2199`), so seeing it while a plan export window is live means the component's control snapshot never saw the window. Two code-verified routes to that, both still live: with `sigenergy_automatic: False`, `automatic_config()` - the only wiring of the component's control entities into Predbat - never runs, so Predbat's plan writes land on its own simulated dummy entities (`inverter.py:539-545`) and a `True == True` read-back is Predbat comparing against itself rather than reading a register; and `fetch_controls` runs on the first tick only (`if first:`, `sigenergy.py:2658`), so a single missed `call_service` event leaves the component's internal state permanently out of step while Predbat's write-skip comparison guarantees the healing write is never re-issued (GH#4894, probe-verified in `test_sigenergy.py`). `_publish_mqtt` logged the live `accessToken` on every command until PR #4926 added `SigenergyAPI.redact()`. The docs template still writes `number.sigen_plant_grid_import_limitation` to 0 in both freeze branches (`docs/inverter-setup.md:1936-1940` and `:1971-1975`), and Sigenergy's own EVAC/EVDC chargers are EMS-throttled to the plant import limit - so the block lands exactly when Predbat asks for Freeze Charging to hold for a car (GH#4911). On the **stock** `has_reserve_soc: False` shape the planned reserve floor is model-only and the battery drains to the inverter's own cut-off — see the "Plan shows a flat hold at the reserve" symptom row (GH#5358). **Every power lever on SIGCLOUD has been silently inert (GH#5376, hardware-verified by the reporter's live MQTT log and battery data; fix in open PR #5377):** `send_battery_command()` writes its six optional power/priority fields — `chargingPower`, `pvPower`, `maxSellPower`, `maxPurchasePower`, `chargePriorityType`, `dischargePriorityType` — at the **top level** of the payload while the required command fields sit inside `commands[0]` (since `cba12cc3`, first released v8.36.11); the API reads the levers from the command objects, so the top-level fields are silently ignored, and planned charge rate, bidirectional export cap, freeze-charge/`freeze_export_hold` `chargingPower: 0` and both priority types (all set in `apply_controls()`, `sigenergy.py`) have never reached the hardware. The sigenergy suite passes while the bug is live because `test_sigenergy.py` asserts `payload["chargingPower"]` at the top level while correctly checking `commands[0]` for the required fields — **a test that locks in the implementation's current output makes suite-green meaningless for that surface**; a fixed test must assert inside `[0]` *and* absence at top level. Suspect, not verified: until such a fix lands and is hardware-tested, the #4761 freeze-export reasoning and the `freeze_export_hold` branch have never acted with real effect, so a post-#5377 regression report about holds/exports is that surface going live for the first time. | `sigenergy` | +| Sigenergy (`sigenergy.py`) | Lifetime/history totals are cumulative server-side and reset around EU midnight, which showed up as an overnight dip; `fetch_history_totals` applies a monotonic clamp. "Energy totals went backwards overnight" starts here. Separately, a user `apps.yaml` override of `inverter.has_reserve_soc: true` (stock `SIG` default is `False`, deliberately — reverted in `5bfc5c80` after #2873/#3124) bypasses two guards that normally make reserve writes inert for Sigenergy (`execute.py:947-950`, `inverter.py:547-549`), so `adjust_reserve()` writes straight to `number.sigen_plant_ess_discharge_cut_off_state_of_charge`. Confirmed from a reporter's log: zero occurrences of the "Inverter does not support reserve" disable message the first guard would normally log, plus 56 reserve writes landing on that register (GH#4728). The stock template's `reserve:` line also points at that same entity rather than `ess_backup_state_of_charge`, which is what makes the override immediately destructive rather than merely inert. `Applying mode=eco ... duration=720min` is emitted only by the no-window fall-through branch (`sigenergy.py:2199`), so seeing it while a plan export window is live means the component's control snapshot never saw the window. Two code-verified routes to that, both still live: with `sigenergy_automatic: False`, `automatic_config()` - the only wiring of the component's control entities into Predbat - never runs, so Predbat's plan writes land on its own simulated dummy entities (`inverter.py:539-545`) and a `True == True` read-back is Predbat comparing against itself rather than reading a register; and `fetch_controls` runs on the first tick only (`if first:`, `sigenergy.py:2658`), so a single missed `call_service` event leaves the component's internal state permanently out of step while Predbat's write-skip comparison guarantees the healing write is never re-issued (GH#4894, probe-verified in `test_sigenergy.py`). `_publish_mqtt` logged the live `accessToken` on every command until PR #4926 added `SigenergyAPI.redact()`. The docs template still writes `number.sigen_plant_grid_import_limitation` to 0 in both freeze branches (`docs/inverter-setup.md:1936-1940` and `:1971-1975`), and Sigenergy's own EVAC/EVDC chargers are EMS-throttled to the plant import limit - so the block lands exactly when Predbat asks for Freeze Charging to hold for a car (GH#4911). On the **stock** `has_reserve_soc: False` shape the planned reserve floor is model-only and the battery drains to the inverter's own cut-off — see the "Plan shows a flat hold at the reserve" symptom row (GH#5358). **Every power lever on SIGCLOUD has been silently inert (GH#5376, hardware-verified by the reporter's live MQTT log and battery data; **fixed in PR #5377, merged 2026-10-03 and first shipped in v9.3.5** — the six optional fields are now built into the command object, `sigenergy.py` — keep the pre-#5377 reading for older logs):** `send_battery_command()` writes its six optional power/priority fields — `chargingPower`, `pvPower`, `maxSellPower`, `maxPurchasePower`, `chargePriorityType`, `dischargePriorityType` — at the **top level** of the payload while the required command fields sit inside `commands[0]` (since `cba12cc3`, first released v8.36.11); the API reads the levers from the command objects, so the top-level fields are silently ignored, and planned charge rate, bidirectional export cap, freeze-charge/`freeze_export_hold` `chargingPower: 0` and both priority types (all set in `apply_controls()`, `sigenergy.py`) have never reached the hardware. The suite passed while the bug was live because `test_sigenergy.py` asserted `payload["chargingPower"]` at the top level while checking `commands[0]` for the required fields — **a test that locks in the implementation's current output makes suite-green meaningless for that surface**; the #5377 fix rewrote the test to assert inside `commands[0]` (its "optional fields that were not passed are omitted" cases are the in-tree example of testing the fixed shape). Suspect, not verified: the fixed surface has never been hardware-tested by the triage bot, so a post-v9.3.5 regression report about holds/exports on SIGCLOUD may be that surface going live for the first time; pre-#5377, the #4761 freeze-export reasoning and the `freeze_export_hold` branch never acted with real effect. | `sigenergy` | | AlphaESS (`alphaess.py`) | Once any discharge window is configured (`ctrDis=1`), AlphaESS firmware allows discharge *only* inside that window — outside it the battery is charge-from-PV-only, with no self-consumption fallback (AlphaESS's own documented behaviour, per GH#4701). Predbat's generic handling of what happens outside a configured discharge window (`execute.py`) has no AlphaESS-specific case and assumes normal demand mode resumes there, which holds for other brands but not this one — the next best export window Predbat schedules can be hours away, leaving the battery locked out of discharge until then (GH#4723). A separate write-ordering bug, **fixed in PR #4776 (merged 2026-08-27)**: the periodic API commits a schedule in stages — window, then enable switch, then target SoC — pressing the write button after each, so the *first* commit of a cycle always carried the *previous* cycle's target while `alphaess_min_write_interval` (default 300s) held the correcting write back for the rest of that pacing window; a manual charge on an already-live slot could run at the stale target for up to five minutes (GH#4769, confirmed from the reporter's log, not inferred). Also: the legacy `/updateDisChargeConfigInfo` endpoint's `batUseCap` field (the reserve/export-SoC value) has no floor on the Predbat side (`alphaess.py:1103`/`:1143`), while the periodic path clamps it to AlphaESS's documented `[10,100]` range — a reserve/export-SoC entity set below 10 gets every legacy write rejected with an undocumented `10001` (GH#4748); that errno is the signature to recognise this class by. **The single most important thing to know here, documented on `main` since PR #4744: the AlphaESS Open API cannot control export at all, so Predbat cannot make it export.** There is no forced-export, working-mode or dispatch endpoint - only a grid-charge window with a target SoC, a discharge window with an SoC floor, or both together on entitled systems. Force Export therefore exports nothing beyond genuine solar surplus, and Freeze Export comes out byte-identical to Demand mode because `gridCharge` gates only *timed grid charging* and nothing can stop the battery charging from solar. The documented answer is to set `select.predbat_mode` to `Control charge` on AlphaESS; left on `Control charge & discharge` Predbat plans exports that never happen *and* programs the discharge window up to `plan_interval_minutes` early, which on permission-window semantics bars the battery from covering house load until the window opens (that is GH#4723's mechanism). Forced export is reachable over local Modbus, never from the cloud API. Separately, PR #4867 changed how a hold is expressed: Predbat now writes a synthetic **enabled 10% target at 100 W** on a stable daily window (`ALPHAESS_HOLD_SOC`/`ALPHAESS_HOLD_POWER`, `alphaess_const.py:183-184`) rather than the old zero-rate hold, because zero power disables charging in Predbat and AlphaESS may discard such a profile. A genuine grid charge is not replaced by it (`alphaess.py:1106`). Note the consequence, raised in review and not tested on hardware: unlike a zero-rate hold this profile is not energy-neutral by construction, so an install sitting below 10% SoC has a standing 100 W charge request while the hold is in force. | `alphaess_control` | | Ohme (`ohme.py`) | A vendored, version-pinned copy of `dan-r/ohmepy` (`ohme.py:33`, `VERSION`). GH#4719 (2026-08-25) found it several minor versions behind upstream with the control routes (max-charge, session rule) pointing at a withdrawn `/v1/chargeSessions/{id}/rule` — a 404 with Spring's `NoResourceFoundException` — while GET/pause/resume/approve kept working since those stayed on v1 upstream. **Fixed on `main` since** (confirmed 2026-08-27): the vendored client now calls the v2 routes upstream moved those endpoints to. A reporter on an older Predbat version seeing control-only failures (GETs/pause/resume fine, target time/percent/preconditioning silently not applying) is hitting this, not a new bug. Separately, and **still live** (GH#4952): `slot_list()` derives slot energy as `watts * hours` (`ohme.py:124`), treating `watts` as a sustained rate when the observed 16A -> 10A -> 6A taper shows it is a per-slot cap - a reporter measured 48.49 kWh modelled against 19.02 kWh of real charge, 2.55x. Ohme publishes `estimatedSoc`, which would give the true delta, and `ohme.py` never reads it anywhere. The same slots are published as dispatches with `location: "AT_HOME"` and no `source` (`ohme.py:670`), so `rate_add_io_slots()` stamps the whole plug-in-to-target block at the low rate. Note the existing tests assert the current `watts * hours` semantics, and `estimatedSoc` is itself extrapolated when the car has no SoC telemetry - so neither side is a drop-in swap. See the car-charging row for why the `predict()` clamp has to be trimmed *after* this, not before. | `ohme` | -| Solcast (`solcast.py`) | `max_kwh` — the array-capacity ceiling used to sanity-check the forecast — is initialised to `9999` (`solcast.py:1339`) and is only ever reassigned on the Forecast.Solar and Open-Meteo branches. On the direct-API and HA-sensor Solcast paths it stays `9999`, so the "Raw forecast exceeds the array ceiling" warning (`solcast.py:1191`) can never fire for Solcast users — confirmed by reading the assignments (GH#4730). That warning is exactly the diagnostic an "impossible PV predicted" report needs, so its absence isn't evidence the ceiling wasn't exceeded. Branch order at `solcast.py:1368` also means a configured `solcast_api_key` always wins over `pv_forecast_*` HA sensors, so comparing Predbat's number against the HA entity compares against the wrong source when both are set. Two more. `solcast_poll_hours: 4.8` is silently truncated to 4: component args come from `COMPONENT_LIST`'s int defaults via `get_arg()`, which does `int(float(value))` (`userinterface.py:266`), and GH#4441's fix lives on the CONFIG_ITEMS route that component args bypass - the reporter's log showed real fetches at exactly 4h and 18x `code 429` against a 10/day hobbyist limit, while `docs/apps-yaml.md` recommends 4.8 for two-array accounts (GH#4925). And `publish_pv_stats()` reads the shared, briefly-rewound `midnight_utc` live (GH#4804) - see "The shared clock is rewound for ~0.8s every hour" above, which has since been confirmed as a real bug class by a separate merged fix. PV calibration self-reference (GH#5116, **fixed in PR #5121** — mechanism kept for older logs): `sensor._pv_forecast_h0` publishes the **calibrated** `power_nowCL` as its state whenever calibration is on, with the raw value only in the `now` attribute, so anything reading the state history gets calibrated output — and calibration itself used to read that history as its "past forecast" while applying its factors to the raw series, which settles the factor at **√(actual/raw)** and corrects only ~half the bias (logged 0.9073x against a measured 0.82). #5121 replaced the read with a dedicated `sensor._pv_forecast_h0_uncalibrated`, published unconditionally and never scaled by `pv_scaling`; `pv_forecast_history()` reads it (scaled by the *current* `pv_scaling`, so a change takes effect over the whole window at once), falls back per-point to the h0 `now` attribute and then the h0 state for points predating the sensor (`now` was added in the same commit that made the state calibrated, so a point without `now` recorded the raw state), and the web chart plots the uncalibrated series. The h0 state-vs-`now` distinction stays live for other consumers. Two general facts from the same PR's review: `history_attribute(attributes=True)` drops any point whose attributes lack the key unless `fallback_to_state=True` (the state path filters `unavailable`/`unknown`; the attribute path only skips them inside the fallback read — `utils.py`), and the DB mirror returns rows with `attributes = {}` on JSON decode failure (`db_engine.py` — the row is still appended), so an empty attribute read on a `db_primary` history does not prove the data predates the attribute. Anything comparing h0 attributes against Predbat's internal PV series must also account for `pv_scaling`, which sits on one side only (`now` is pre-scaling, the internal series post-scaling). Test-fidelity trap: a calibration test that patches `history_attribute_to_minute_data` cannot see which series calibration learned from — the real path needs a `get_history_wrapper` mock with HA-shaped entries wrapped in `patch_now_utc_exact()`, because `now_utc_exact` is real wall-clock while `minutes_now`/`midnight_utc` are frozen; `prune_today`'s default `group=15` also drops history points closer than 15 minutes apart. `pv_calibration()`'s past-forecast history additionally flows through the **module-global** `history_attribute_to_minute_data` (solcast.py imports it by name), so an instance-level override cannot intercept it — patch `solcast.history_attribute_to_minute_data` and restore it in a `finally`, and re-patch before a second call in the same test. **Open-Meteo shading is client-side (GH#5224):** the Open-Meteo `v1/forecast` API has no `horizon` parameter (closest solar-geometry options are `tilt`/`azimuth` for GTI), so any horizon feature — Predbat's or a user asking for one — must be applied client-side to the downloaded irradiance, which is exactly what the HA Open-Meteo Solar Forecast integration (rany2) does; Predbat's existing knob is the static per-array monthly `shading_factors` applied in `gti_hourly_to_period_kwh()` (`solar_model.py`), which cannot express geometric blocking. *Suspected, not verified:* that integration's published attribute names could not be confirmed compatible with Predbat's HA-sensor PV path (`detailedForecast`/`forecast` with `pv_estimate` entries), so do not recommend pointing Predbat at its sensors as a workaround without checking the entity shape. **Band built from the wrong series with calibration off (GH#5345, fixed in PR #5347 \[v9.3.4\]: `pv_calibration()` now takes `calibrate_band` and builds the band on the P50 it is returned with — the switch-off path swaps both p50 and the band back to raw, and Open-Meteo keeps its own ensemble band instead of a synthesised `create_pv10` one; keep the pre-#5347 mechanism for older logs):** on a `create_pv10` source (Open-Meteo, Forecast.Solar) with `metric_pv_calibration_enable` off, `pv_calibration()` built `pv_forecast_minute10`/`90` and the published `pv_estimate10`/`90` from the **calibrated** series and only swapped P50 back to raw at its final `return`, so P10 could exceed P50 and P90 fall below it, and `get_cloud_factor`'s (ΣP50 − ΣP10) gap was arbitrary - reported as "the saw tooth has gone". Signature in `sensor._pv_today` `detailedForecast`: `pv_estimate10 / pv_estimateCL` and `pv_estimate90 / pv_estimateCL` constant in every slot while `pv_estimateCL / pv_estimate` varies. Dead end: it is not Open-Meteo's ensemble P10 being "confident" - at the time that download was fetched and then overwritten by `pv_calibration()`, because every Open-Meteo branch set `create_pv10`; the same change makes `fetch_pv_forecast()` keep an ensemble band (`create_pv10` False, `calibrate_band` True) and fall back to the history-based band only when the ensemble download fails. The band is the ensemble's P10 and P90 **as ratios of the ensemble's own median**, applied to the deterministic P50, not the ensemble's absolute values: the P50 comes from the best-match deterministic call and the spread from `icon_seamless`, two different model runs that can disagree on a whole day's level - live on 2 Oct 2026 the absolute ensemble P10 sat above the deterministic P50 for a day three ahead (clamped, leaving a 2% downside) and at 0.28 of it elsewhere. Known weakness of the ratio: the deterministic forecast is re-issued between ensemble runs and moved one site's next-day total from 1635 to 3887 Wh/m2 in ten minutes, above the ensemble's own P90, and P50 x (P90/median) then overshoots anything the sky can deliver - only the `1.2 x kWp` power ceiling bounds it, there is no clear-sky cap. So an implausibly high `pv_estimate90` on an Open-Meteo site maps here. Test trap from that work: the solar test mock returns the first response whose key is a substring of the URL, and `api.open-meteo.com` is a substring of `ensemble-api.open-meteo.com`, so a test registering both in that order silently serves the forecast body to the ensemble call and exercises the no-ensemble path. The in-function `enabled_calibration` flag means "enough history", not the user switch, so the fixed 0.7/1.3 band does not apply with the switch off. Routine SolarAPI Info lines are not in shipped logs; read the `pv_today`/`pv_forecast_h0` attributes instead. | `solcast` | +| Solcast (`solcast.py`) | `max_kwh` — the array-capacity ceiling used to sanity-check the forecast — is initialised to `9999` (`solcast.py:1339`) and is only ever reassigned on the Forecast.Solar and Open-Meteo branches. On the direct-API and HA-sensor Solcast paths it stays `9999`, so the "Raw forecast exceeds the array ceiling" warning (`solcast.py:1191`) can never fire for Solcast users — confirmed by reading the assignments (GH#4730). That warning is exactly the diagnostic an "impossible PV predicted" report needs, so its absence isn't evidence the ceiling wasn't exceeded. Branch order at `solcast.py:1368` also means a configured `solcast_api_key` always wins over `pv_forecast_*` HA sensors, so comparing Predbat's number against the HA entity compares against the wrong source when both are set. Two more. `solcast_poll_hours: 4.8` is silently truncated to 4: component args come from `COMPONENT_LIST`'s int defaults via `get_arg()`, which does `int(float(value))` (`userinterface.py:266`), and GH#4441's fix lives on the CONFIG_ITEMS route that component args bypass - the reporter's log showed real fetches at exactly 4h and 18x `code 429` against a 10/day hobbyist limit, while `docs/apps-yaml.md` recommends 4.8 for two-array accounts (GH#4925). And `publish_pv_stats()` reads the shared, briefly-rewound `midnight_utc` live (GH#4804) - see "The shared clock is rewound for ~0.8s every hour" above, which has since been confirmed as a real bug class by a separate merged fix. PV calibration self-reference (GH#5116, **fixed in PR #5121** — mechanism kept for older logs): `sensor._pv_forecast_h0` publishes the **calibrated** `power_nowCL` as its state whenever calibration is on, with the raw value only in the `now` attribute, so anything reading the state history gets calibrated output — and calibration itself used to read that history as its "past forecast" while applying its factors to the raw series, which settles the factor at **√(actual/raw)** and corrects only ~half the bias (logged 0.9073x against a measured 0.82). #5121 replaced the read with a dedicated `sensor._pv_forecast_h0_uncalibrated`, published unconditionally and never scaled by `pv_scaling`; `pv_forecast_history()` reads it (scaled by the *current* `pv_scaling`, so a change takes effect over the whole window at once), falls back per-point to the h0 `now` attribute and then the h0 state for points predating the sensor (`now` was added in the same commit that made the state calibrated, so a point without `now` recorded the raw state), and the web chart plots the uncalibrated series. The h0 state-vs-`now` distinction stays live for other consumers. Two general facts from the same PR's review: `history_attribute(attributes=True)` drops any point whose attributes lack the key unless `fallback_to_state=True` (the state path filters `unavailable`/`unknown`; the attribute path only skips them inside the fallback read — `utils.py`), and the DB mirror returns rows with `attributes = {}` on JSON decode failure (`db_engine.py` — the row is still appended), so an empty attribute read on a `db_primary` history does not prove the data predates the attribute. Anything comparing h0 attributes against Predbat's internal PV series must also account for `pv_scaling`, which sits on one side only (`now` is pre-scaling, the internal series post-scaling). Test-fidelity trap: a calibration test that patches `history_attribute_to_minute_data` cannot see which series calibration learned from — the real path needs a `get_history_wrapper` mock with HA-shaped entries wrapped in `patch_now_utc_exact()`, because `now_utc_exact` is real wall-clock while `minutes_now`/`midnight_utc` are frozen; `prune_today`'s default `group=15` also drops history points closer than 15 minutes apart. `pv_calibration()`'s past-forecast history additionally flows through the **module-global** `history_attribute_to_minute_data` (solcast.py imports it by name), so an instance-level override cannot intercept it — patch `solcast.history_attribute_to_minute_data` and restore it in a `finally`, and re-patch before a second call in the same test. **Open-Meteo shading is client-side (GH#5224):** the Open-Meteo `v1/forecast` API has no `horizon` parameter (closest solar-geometry options are `tilt`/`azimuth` for GTI), so any horizon feature — Predbat's or a user asking for one — must be applied client-side to the downloaded irradiance, which is exactly what the HA Open-Meteo Solar Forecast integration (rany2) does; Predbat's existing knob is the static per-array monthly `shading_factors` applied in `gti_hourly_to_period_kwh()` (`solar_model.py`), which cannot express geometric blocking. *Suspected, not verified:* that integration's published attribute names could not be confirmed compatible with Predbat's HA-sensor PV path (`detailedForecast`/`forecast` with `pv_estimate` entries), so do not recommend pointing Predbat at its sensors as a workaround without checking the entity shape. **Band built from the wrong series with calibration off (GH#5345, fixed in PR #5347 \[v9.3.4\]: `pv_calibration()` now takes `calibrate_band` and builds the band on the P50 it is returned with — the switch-off path swaps both p50 and the band back to raw, and Open-Meteo keeps its own ensemble band instead of a synthesised `create_pv10` one; keep the pre-#5347 mechanism for older logs):** on a `create_pv10` source (Open-Meteo, Forecast.Solar) with `metric_pv_calibration_enable` off, `pv_calibration()` built `pv_forecast_minute10`/`90` and the published `pv_estimate10`/`90` from the **calibrated** series and only swapped P50 back to raw at its final `return`, so P10 could exceed P50 and P90 fall below it, and `get_cloud_factor`'s (ΣP50 − ΣP10) gap was arbitrary - reported as "the saw tooth has gone". Signature in `sensor._pv_today` `detailedForecast`: `pv_estimate10 / pv_estimateCL` and `pv_estimate90 / pv_estimateCL` constant in every slot while `pv_estimateCL / pv_estimate` varies. Dead end: it is not Open-Meteo's ensemble P10 being "confident" - at the time that download was fetched and then overwritten by `pv_calibration()`, because every Open-Meteo branch set `create_pv10`; the same change makes `fetch_pv_forecast()` keep an ensemble band (`create_pv10` False, `calibrate_band` True) and fall back to the history-based band only when the ensemble download fails. The band is the ensemble's P10 and P90 **as ratios of the ensemble's own median**, applied to the deterministic P50, not the ensemble's absolute values: the P50 comes from the best-match deterministic call and the spread from `icon_seamless`, two different model runs that can disagree on a whole day's level - live on 2 Oct 2026 the absolute ensemble P10 sat above the deterministic P50 for a day three ahead (clamped, leaving a 2% downside) and at 0.28 of it elsewhere. Known weakness of the ratio: the deterministic forecast is re-issued between ensemble runs and moved one site's next-day total from 1635 to 3887 Wh/m2 in ten minutes, above the ensemble's own P90, and P50 x (P90/median) then overshoots anything the sky can deliver - only the `1.2 x kWp` power ceiling bounds it, there is no clear-sky cap. So an implausibly high `pv_estimate90` on an Open-Meteo site maps here. Test trap from that work: the solar test mock returns the first response whose key is a substring of the URL, and `api.open-meteo.com` is a substring of `ensemble-api.open-meteo.com`, so a test registering both in that order silently serves the forecast body to the ensemble call and exercises the no-ensemble path. The in-function `enabled_calibration` flag means "enough history", not the user switch, so the fixed 0.7/1.3 band does not apply with the switch off. Routine SolarAPI Info lines are not in shipped logs; read the `pv_today`/`pv_forecast_h0` attributes instead. A source that publishes a 0-PV day is not a safe fallback either — the counter-argument to any "just publish 0" suggestion: a 0-yield day is a `pv_calibration()` down-day (actual < 10% of forecast) and ≤ 2 valid days switches calibration off entirely, leaving the raw forecast (GH#5388 context, code-read on main; `solcast.py`). | `solcast` | | HA write/verify (`ha.py`, `inverter.py`) | `get_state(refresh=True)` is a no-op in a normal HA add-on install: it only re-reads when `not self.ha_key`, and `ha_key` is set in that case, so verification reads the websocket cache, never a fresh value (`ha.py:798`, since PR #2342). Separately, `write_and_poll_value()` treats `domain == "sensor"` as Predbat-owned and POSTs `/api/states/` directly (`inverter.py:2155`, `ha.py:1071-1077`) instead of calling the integration. Probed empirically across all four write-path/domain combinations (GH#4738): a control entity that resolves to a `sensor.*` domain by mistake gets a direct state overwrite that reads back as success, with zero corresponding lines in the real integration's log. `call_service_wrapper()`'s return value used to be discarded at every `inverter.py` write site, so a rejected HA service call (e.g. a Modbus "Illegal Function" error) never reached the verify logic. That is no longer true everywhere: PR #4878 made `call_service_template()` check it (`inverter.py:2817`) and log "was not accepted, it will be retried next cycle". The other eleven `call_service_wrapper()` call sites in that file still discard it, so check the specific write path before repeating the general claim. Later triage extended this row rather than replacing it. `write_and_poll_switch()` builds its service name as `domain + "/turn_..."` straight from the entity's own domain with no validation (`inverter.py:2102`), and `adjust_inverter_mode()` routes every `has_ge_eco_toggle` inverter through it - so an `inverter_mode` pointed at a `select.*` entity produces `select.turn_on`, an HA-invalid service retried 10x at 10s per attempt, and the string coercion just below (`inverter.py:2110-2111`) maps any select state to False, so the force-export direction reports success having written nothing (GH#4909, still live — though on 3-phase GEC the real problem is one stage deeper: no eco register exists at all, and the write path is skipped with `No entity_id for ECO Toggle`; see the GE Cloud row). `adjust_force_export`/`adjust_charge_window` format their times at plan precision with no snapping (`TIME_FORMAT_HMS`, `inverter.py:2608`/`:2615`), so against an integration whose select only offers 15-minute options (Solar Assistant/Deye) Predbat writes e.g. `20:55` and `write_and_poll_option` retries the impossible value 10x every cycle; nothing anywhere reads a select's `options` attribute, and `charge_time_entity_is_option` only decides *how* an entity is written, never which values are legal. The fix precedent is already in tree - AlphaESS's `snap_time_grid()` (`alphaess_const.py:215`), snap start forward and end backward with a collapse check (GH#4947). A second, independent way plan-precision times fail on the same select path: `write_and_poll_option()` only strips the seconds from an 8-char `HH:MM:SS` write when the *previous* state is a healthy 5-char `HH:MM:SS` read — the read feeding it returns `None` for an entity missing from the state cache (`ha.py:817`) or the literal `"unknown"`/`"unavailable"` when degraded (`ha.py:815`), either skips the strip and the raw `HH:MM:SS` goes to `select_option`, which HA rejects; all 10 retries re-read and fail identically, and the log reads `didn't complete got unknown` (GH#5009, probe-verified on main). That degraded-read shape can also be **permanent**: an SA slot select can report `unknown` for days — every write read-back `unknown`, zero verified writes, across an SA rebuild and both native-API and MQTT transports — so both symptoms (read path `unable to read Export window - ... returned no data` every cycle via `time_string_to_stamp("unknown")` → `None`, and the write path above) fire continuously; state reporting varies by SA backend (GH#5214). **A user edit of `config.py`'s INVERTER_DEF `charge_time_format` to anything ≠ exactly `"HH:MM:SS"` is a masking workaround, not a fix (GH#5214):** the `!= "HH:MM:SS"` branch (`inverter.py`) swaps apps.yaml entities for Predbat's own `sensor.predbat_{type}_{id}_*` dummies — reads parse Predbat's own echo and writes become direct `set_state` POSTs that never reach the inverter — so the log goes quiet while SA window control is silently severed. Tells, all observed in #5214's attachments: a `Creating dummy entity sensor.predbat_{TYPE}_{id}_charge_start_time` line in the log (`create_entity()` only logs when the entity is missing from HA state, so its presence proves the ≠ branch ran); the debug yaml's live `args` naming dummy ids while `args_from_apps_yaml` still lists the real `select.*` entities; and warnings stopping after a user restart then resuming after an addon update (which overwrites the edited `config.py`, wiping it). Also: pre-#3533 SA shipped `charge_time_format: "S"` (dummy time entities, no SA select wiring), so "SA time warnings started after updating" can mean the user crossed the #3533 template wiring rather than a format regression. The fix that serves both is reading the select's `options` attribute — nothing in the codebase does today. Six of the eight `adjust_*` writers discard the write helper's boolean and fire `call_notify()`/`mqtt_message()` regardless, and `rest_setReserve()` has the same gap (GH#4845). `call_service_template()` used to record its dedup hash *before* the call, so a silently failed `select_option` was skipped as "previously called" on every later cycle; fixed (it now records only on success and drops the record on failure), though it still deliberately returns True so callers cannot downgrade a freeze into a stop (GH#4876). From that same report, a false-alarm class worth recognising: a `didn't complete got X` where X is exactly the *previous* cycle's target, on an integration that republishes its entities slowly, is the verify window reading HA's cache - the write had already landed by the next cycle. That window is arithmetic, not a constant: up to `INVERTER_MAX_RETRY` (10) writes, each polled through one `inv_write_and_poll_sleep` budget with a doubling read interval - **~40s on GS, ~100s on GE/GEC/GEE** (`write_and_poll_sleep` is 10 for those types, 4 for GS). **Per-type since PR #5297 (merged 2026-09-29):** `INVERTER_DEF[type]["write_max_retry"]`/`["write_backoff"]` (consumed by `_write_attempts()`/`_write_backoff_result()`), and currently only `GWMQTT` deviates - 3 attempts, 10s apart - plus a per-control degraded state for those types: a write of the *same* target failing `INVERTER_WRITE_BACKOFF_FAILURES` (2) times in a row drops to one attempt per `INVERTER_WRITE_DEGRADED_INTERVAL` (300s) until a write verifies, the read-back matches or a different target arrives (full ladder restored at once); a degraded control still reads back, logs and records the failure every call - nothing is skipped silently. Every other type keeps 10 attempts and its existing sleep. A settling bounce that outlasts it is a second false-alarm shape (GH#5130): a HA `number.*` entity can step through intermediates (0 → a float → the settled int) for tens of seconds after a timed-slot write (~44s observed on solax_modbus), so a `didn't complete got ` whose value sits *between* the old and target values is a settling read-back, not a failed write - it self-heals next cycle when the initial read matches. The bounce reads feed only `value_matched()` (keep polling vs rewrite) and only the final read reaches the control ledger (`record_write` on match, `clear` on failure), so the worst case is a redundant same-value rewrite plus the false warning; the in-repo precedent for re-reading after settling is Solis Cloud's `verify_settle_seconds` re-read. And the discharge-stop fallback (GH#5094, still live): `adjust_charge_immediate()` issues a "stop discharge" service step before the charge/freeze step, and when `discharge_stop_service` is not configured it falls back to `charge_stop_service` with `domain="discharge"` - a documented fallback dating to #1614. On Tesla service-hook setups `charge_stop_service` maps to `self_consumption` + backup reserve 0, which on a Powerwall is **active discharge to load** (the EV charger counts as load), not a stop - so every freeze/hold cycle Predbat sends "discharge normally" and then re-commands "hold" ~2s later, and Tesla occasionally applies the first and drops the second, so the battery discharges into the car until the next 5-minute cycle. Verified from the reporter's log: the exact pair fires from the carHolding branch ("Disabling battery discharge whilst car 0 is charging" → `adjust_charge_immediate(soc_percent, freeze=True)`); the reporter had commented out `discharge_stop_service` (and `discharge_start_service`) when disabling export control, which is what exposed the fallback. Workaround: define `discharge_stop_service` in apps.yaml with commands that actually stop discharge (for a no-export setup, the same commands as `charge_freeze_service`). Caveat on the reporter's own `repeat: False` idea: enabling the service dedup means accepted calls are not re-sent (see the GH#4876 entry above - HA accepts the call but Tesla firmware can still drop it), so dropped freeze commands would never be retried. Both this and the fixed export clobber below stem from the same property of the Tesla template set: `charge_stop_service` and the freeze/discharge templates all write the same `operation_mode` select, and the templates mean opposite things on this hardware. **A sibling of the same family, over the reserve register rather than the mode selects (GH#5300, triaged 2026-09-29 against d951201e, reporter on a hand-built apps.yaml):** inside one 5-minute pass the hook charge (`adjust_charge_immediate()` issues `charge_freeze_service`/`charge_start_service` via `call_service_template`, inverter.py) and `execute.py`'s reserve reset (`if self.set_reserve_enable and resetReserve: inverter.adjust_reserve(0)`) both touch the Powerwall's reserve and the reset has **no isCharging guard** - `resetReserve` initialises True whenever `set_charge_window or set_export_window` and is cleared only by the freeze / hold-charging / car / iBoost hold branches - while `adjust_reserve()` floors at `reserve_percent` and writes through `write_and_poll_value()`, so the reset always ends up holding the register, against hook writes that are fire-and-forget. A service hook that uses the reserve register as its charge signal loses that race on every charge-window cycle. The shipped template is immune by two mechanisms, both worth checking before triaging any "Tesla reserve overwritten" report: `templates/tesla_powerwall.yaml` sets `has_reserve_soc: False` in the `inverter:` block (`fetch_inverter_data()` then force-disables `set_reserve_enable`/`set_reserve_hold`, logging `Note: Inverter does not support reserve - disabling reserve functions`), and with no `reserve:` arg wired Predbat dummies the arg (`create_missing_arg`/`create_entity`, inverter.py) so `adjust_reserve` writes land on a Predbat placeholder - the Sigenergy dummy-arg shape again. Discriminator for real-vs-dummy reserve wiring: the `Inverter N Current Reserve is X% ... and new target is Y%` and `Wrote X to reserve, successfully now Y` lines name the register, not the entity - a non-zero current read means the `reserve:` arg resolves to a real wired entity; a dummy reads 0 and at an equal target prints `already at target` instead. Dump-config tells from the same report: a `set_reserve_enable` switch showing `value: null` in a debug dump can still resolve True (expert-mode-switched item, hidden and entity not created while expert mode is off), and unknown keys in the dump's `args:` mean a hand-built config - do not assume the shipped template. | `inverter` | | GE Cloud (`gecloud.py`) | Gateway fields can come back null. `merge_non_null` stops nulls overwriting good values, and the publish path guards null containers — GH#4656 was a null crash in that area. `number_event` clamps an out-of-range write silently and does not report the clamped value back, while `write_and_poll_value()` keeps comparing against the original pre-clamp target - a permanent false `had_errors` and an unbounded retry (GH#4826). PR #4828 clamped inside `adjust_reserve()` only; the shared write helper still has no min/max awareness and the other `inverter.py` write sites still pass unclamped targets (GH#4831). The error-code path has since been reworked (`9560c61a`, `d6e5998c`) after GH#4896 showed `success: false` nulled `data` before the code branch could read it, so codes GivEnergy documents as never succeeding on a repeat (-3/-4/-7) were retried the full budget and `message` was never surfaced. Separately, **fixed in PR #4954 (merged 2026-09-06)** - keep the mechanism for anyone holding an older log: for `GE`/`GEC`/`GEE` the max battery rate was taken solely from the `charge_rate` entity's `max` attribute, and the `elif "battery_rate_max"` branch below it was unreachable for those types. Percentage-rated models (the 3-phase units) have no absolute charge power register, so `ge_cloud_automatic` left `charge_rate` unset and the attribute read returned the 2600 W default; both the correct rate GECloud writes into `battery_rate_max` and the user's own apps.yaml value went unread, and every rate was clamped to 2600 W. On those models it also corrupted the writes, since the percentage registers are set as `new_rate / battery_rate_max_raw * 100` against a denominator 3.8x too small. The fix falls back to `battery_rate_max` only when `charge_rate` resolves to nothing (GH#4908). GH#4198 is the same code shape on SolaX and is not covered by that fix. A model-string parser gap sits above that path (GH#5136): **the integer-rating half was fixed in PR #5293 (merged 2026-09-30, v9.3.3)** — `get_max_inverter_rate_from_model()` now parses an integer segment straight after the `3HY` token (`GIV-3HY-11` → 11 kW) *before* the decimal pass, so a later decimal segment cannot hide it; it also guards a null model, strips trailing punctuation (`GIV-HY-8.0-G3-HV.`) and warns readably on an unparseable rating ("using max charge rate instead - set inverter_limit in apps.yaml"); keep the pre-#5293 signature (decimal-only match, so an integer-rated `GIV-3HY-11` fell back to `max_charge_rate` — which the GE Cloud API can change independently, 11000 → 9984 W observed). **Still live:** `publish_info()` reads only `max_charge_rate`, never the `max_discharge_rate` in the same payload, and `battery_rate_max` is bound to the `_max_charge_rate` entity — a model whose discharge rating differs keeps the wrong rate. Distinct from GH#4908 (the absent-`charge_rate` 2600 W default). Start on any "GEC inverter rate is wrong" report by checking the model string for a decimal point; the tests now cover the integer-3HY cases too. `automatic_config()` binds `inverter_limit` to the resulting `_max_inverter_rate` sensors and `prediction.py` clips PV against `inverter_limit`, so the error reaches the plan. **Mixed fleets: the rate read/write gates test key existence, not per-slot state (GH#5359, probe-verified on main b161949d via `tools/triage_test.sh`):** `get_current_charge_rate()`/`get_current_discharge_rate()` (`inverter.py`, the `"charge_rate_percent" in self.base.args` branch) and `adjust_charge_rate()`/`adjust_discharge_rate()` (which write **both** keys, each gated only on key existence) interrogate whether the key exists **anywhere in the fleet's args**, not whether *this* inverter's slot is set — and `get_arg()` on a `None` slot applies the caller's default quietly (`userinterface.py`, float branch) — so in a fleet mixing GEC percent-controlled units (3-phase `GIV-3HY` has no `battery_charge_power` watts register; `automatic_config()` then sets `charge_rate_percent` and `build_entities()` leaves per-device `None` slots, deleting `charge_rate` when no device has the register) with watts-controlled units, a percent slot that is `None` reads quietly as 100% → `battery_rate_max_raw`, and every rate write for the power-mode sibling lands on a missing entity (`Warn: ... No entity_id for charge_rate to write ...` + `record_status(had_errors=True)`, so "Read-Only with Errors" too). The low-power rate search, which steps down relative to the read-back, then restarts from the maximum every cycle and never converges — #3311's fault extended to mixed fleets (the single-inverter shape is in the low-power symptom row). PR #4645 (open, for #3311) creates its per-slot rate dummies but does not touch these four gates — the #4908 battery-sizing read (`get_arg("charge_rate", indirect=False, index=self.id)` before falling back, inverter.py:583) is the per-entry pattern the gates lack. One API-shape trap on the site endpoint (`GE_API_SITE`, `site/{uuid}`, added for PR #4978): the published OpenAPI spec types `limits.import`/`limits.export` as a string with a null example and never shows a populated one, so the shape had to be learned from live accounts. A populated limit is `{'enabled': bool, 'power': {'watts': int, 'amps': float}}`, and **`enabled` is the thing that matters** - a site describes its connection's declared capacity in exactly the same structure as a curtailment, marked disabled, so `power.watts` alone tells you nothing about whether anything is being restricted. Site 64744 gave `{'import': {'enabled': False, 'power': {'watts': 46000, 'amps': 200}}, 'export': {'enabled': True, 'power': {'watts': 4500, 'amps': 19.6}}}` - a 200A supply whose export really is curtailed to 4.5kW - while site 46714's disabled 6kW export is its declared capacity and no restriction at all. A fleet survey confirmed the flag across accounts, so only an enabled limit is applied; an enabled zero is a real zero-export connection. A site with no limit reports a null import/export. For a fleet survey use `GET /v1/site`, the list endpoint, which returns every site's `limits` and the inverters under it (with `info.max_discharge_rate`) in one response - no second call per site. Two GEC triage tools and one clobber trap (GH#5017). The component publishes an entity for every `/settings` entry with no name filter (`async_get_inverter_settings()`/`publish_registers()`, `regname_to_ha()` is plain lowercasing), so **absence of an entity in the reporter's debug yaml is proof the API never listed the register** — "the app can control it" does not imply the settings API exposes it, because the app/GivTCP write via local control paths rather than the `/settings` list (a GIV-3HY-20 had slot 1/2 lower-SoC registers absent while 3..10 existed). Caveat: the settings list is fetched once per process start (`register_list` cache) and only the parsed flags are logged, so the debug yaml is the only observable. To capture the raw register list, `python3 apps/predbat/gecloud.py --api-key --write-entity number.probe_nothing 0` makes `test_gecloud_direct()` print the full "Available entities/registers" list with raw cloud names and setting IDs and performs no write — but the harness's startup pass also runs `enable_default_options()` against the real inverter (floors reset, unused windows zeroed), so say so when asking a user to run it. Clobber trap: the missing-feature branch `set_arg("discharge_target_soc", None)` *deletes* the key, including an explicit apps.yaml entry the user wrote; `set_arg_auto(..., overwrite=False)` (component_base.py) is the in-repo pattern for letting an explicit value win — check the `has_*` detection branch on any "automatic config deleted my setting" report. The `inverter_mode` half of this class was fixed in PR #5059 (`885c637d`) — it now binds via `set_arg_auto(..., overwrite=False)` like `export_limit` — but the other missing-feature deletes are still plain `set_arg(..., None)` on main: `pause_mode`, `pause_start_time`, `pause_end_time`, `discharge_target_soc`, `charge_rate_percent`, `discharge_rate_percent`, `givtcp_rest`. A companion test-fidelity trap (GH#5056, verified on main): the component test mocks (`MockBase.set_arg` in `test_ge_cloud.py`; same shape in `test_ohme.py` and `test_alphaess_api.py`, whose mock `set_arg_auto` also writes unconditionally) store the value instead of mirroring the real `set_arg()`'s None-deletes, and their `args_from_apps_yaml` is empty so `set_arg_auto()` falls through to the plain path — so the "feature not detected" assertions (`config_args.get(...) is None`) pass identically whether the code stored None or deleted the key, and a green suite cannot distinguish fixed from unfixed. Any test for this class must populate `args_from_apps_yaml` and assert the manual value survives — #5059's Test 10 in `test_ge_cloud.py` is now the in-repo precedent. And the deeper half of GH#4909 (still live): the 3-phase GEC class exposes **no** eco/demand/mode register in `inverter/{sn}/settings` — verified by enumerating the `predbat_gecloud_*` entities in the reporter's debug dump — yet the `GEC` def sets `has_ge_eco_toggle: True` unconditionally (`config.py`) and auto-config maps `inverter_mode` via `build_entities("switch", ["enable_eco_mode"])` (gecloud.py), which emits nothing when the register is missing, so `adjust_inverter_mode()` logs `No entity_id for ECO Toggle` and returns without writing every plan cycle and freeze export silently no-ops (force-discharge paths work fine). Fix wrinkle for the maintainer: the flag is one per-inverter-type bool while `build_entities` works per-device. The freeze-export half of the same class (GH#5082, follow-up to GH#4909, code-verified on main): `inv_has_timed_pause` is probed at runtime (inverter.py flips it False when the `pause_mode` entity is absent/unavailable) but `inv_support_discharge_freeze` is read only from the static `INVERTER_DEF` entry (`config.py` sets it True for GEC), so `execute.py` never force-disables `set_export_freeze` for the eco-less/pause-less 3-phase class and a freeze-export plan can never be executed. Check both flags on any "freeze silently didn't work on GEC" report - and do not grep only for `No entity_id for ECO Toggle`: that warning comes from the eco branch, while freeze export on this class fails primarily through the pause/rate branch, so the eco warning being absent says nothing about the freeze path. Fix direction: mirror the `inv_has_timed_pause` presence-probe; the manual-freeze-export override path already degrades correctly when `set_export_freeze` is False (plan.py drops to demand). Hardware-reported on the GIV-3HY class: `charge_power_rate` limits AC (grid) charging only - `battery_power` stays at the full discharge figure while the charge register reads 1% - so the `adjust_charge_rate(0)` fallback in the freeze branch is a no-op for PV charging there; note `charge_discharge_with_rate` is False for GEC, so the first `adjust_charge_rate(0)` in each branch is skipped and the write only happens via the `has_timed_pause`-else fallback. Fix precedent for building a freeze from registers the device actually has: `inv_support_feedin_first` + Fox Cloud freeze export (PR #5038); the GIV-3HY exposes `enable_force_discharge` + `discharge_down_to_percent` + `battery_reserve_percent`, the raw material for the #4909/#5082 option-2 freeze. And since #5056 a manual `inverter_mode` apps.yaml override is silently deleted by automatic config, so "just map the eco toggle manually" is not currently viable either. The neighbouring **charge-side** gap (GH#5040: `scheduled_charge_enable` had no `enable_force_charge` candidate, so a 3-phase GEC that gates timed grid charge behind both charge switches logged real charge states while importing nothing) is **fixed in PR #5042** (`a7121135`, v9.0.2): auto-config now binds `scheduled_charge_enable` to `enable_force_charge` with `ac_charge_enable`/`enable_ac_charge` as fallbacks, and `enable_default_options()` holds `enable_ac_charge` on as a static enable when the force register exists. Keep the mechanism for pre-v9.0.2 logs, where the signature is a 3-phase GEC showing plan charge states against flat grid power with no warning anywhere. The #5042 gate itself had a second gap (**GH#5269, fixed in PR #5270, merged 2026-09-27** — keep the pre-#5270 signature): `enable_ac_charge` was written in exactly one place, `enable_default_options()` — first cycle after start-up, then once per 24h (`default_options_stamp`) — while Predbat's own `scheduled_charge_enable` write (`switch_event()`) checked nothing about the gate, so any external writer (an Axle VPP clean-up, the GE app, another integration) could clear `enable_ac_charge` and Predbat re-asserted it only on that 24h pass or a restart; "restarted and it worked" is part of the pre-fix signature, as is `enable_force_charge` on with `Charging target …%` for hours, SoC flat and no error. #5270 re-asserts the gate from `switch_event()` when a force-charge write succeeds (`ensure_ac_charge_gate`, with Axle's clean-up read as read-only for the 24h pass and a binding guard), so the "gate cleared between cache refreshes" residual window is the remaining maintainer call on that fix. So "3-phase GEC plans a charge, status says Charging, battery imports nothing" now has **two** signatures: pre-v9.0.2 = never bound (#5040), v9.0.2+ pre-#5270 = externally cleared and not re-asserted (#5269); on a post-#5270 system check whether the force switch Predbat drove is actually the `enable_force_charge` register. GE Cloud EMS plants (GH#5103, **fixed in PR #5109, merged 2026-09-24** — keep the mechanism for pre-#5109 logs): with an EMS discovered `polling_mode` is set False (`gecloud.py`), and the settings refresh used to re-read each battery inverter's `/settings` only on the first tick of the process — so the AC3 registers, including the DC-discharge slot windows, were read once and published stale for the rest of the run, while live status/meter polling kept looking normal. #5109 re-reads settings hourly (`SETTINGS_SLOW_REFRESH_SECONDS`, `gecloud.py`), refreshes the EMS and gateway devices every cycle, and runs `check_ems_inverter_slots()` on every snapshot — which now *warns* when a battery inverter's slot 1 is not the full-day window (the #3781 rule: `Inverter X slot 1 is not 00:00-23:59 ... it will override the EMS and can stop charging or discharging early`), once per episode — recorded in `ems_slot_warned` so a standing misconfiguration does not inflate the error counter on every refresh, so the non-00:00–23:59 slot 1 fault is a visible warning on current main and silent only pre-#5109 (pre-fix signature: plan/status look right but the batteries hold 0 W past a fixed wall-clock time every night). Under EMS auto-config Predbat still writes mode/schedule-enable/rates to the AC3s. Nuance: the startup read itself can still be skipped when the storage settings cache is under 10 minutes old (`settings_from_cache`), so even a restart is not guaranteed to re-snapshot the registers. **The option-validation parser used to crash-loop the whole component (GH#5217, fixed in PR #5218, merged 2026-09-26 — keep the signature for pre-#5218 logs):** `select_event()` and `publish_registers()` both `split("(")` a cloud `Value must be one of: (...)` string, so an option label containing `(` raised `ValueError: too many values to unpack (expected 2)` — the component does not die (`ComponentBase.start()` catches it), it *crash-loops*: `api_started` never set, automatic config never runs, no plan, `Error: GECloudDirect: too many values to unpack (expected 2)` every cycle. #5218 added the module-level `parse_validation_options()` (`gecloud.py`): split on the first `(` and strip only the final `)` (so inner brackets in labels survive), and `None` for a string with no `(` triggers a warning and falls back to the setting's `in:` validation-rule options instead of raising — the paren-less guard is required, not optional, because `split("(", 1)` alone fails the other direction. Which GE Cloud setting actually returns a multi-`(` validation string was never captured (the storage cache saved at the settings write would be the fixture to ask for). **`refresh_discovery()` is not device re-discovery, and GE Cloud re-discovers hourly since PR #5295 (GH#5294, merged 2026-09-29 - the symbols below were read against the pre-fix tree):** `ComponentBase.refresh_discovery()` only rebuilds the capability-catalogue *report* from data already in hand (a component opts in via `build_discovery()`) and never calls the component's API - the *similarly named* `refresh_devices()` is the thing that calls it. Pre-#5295 that device discovery (`async_get_devices()`, EV chargers, EMS/gateway selection, one-shot automatic config) ran exactly once inside `if first:` and `first` flips off permanently after the first successful run, so on any "GE Cloud didn't pick up my new inverter/charger" or "removed inverter still shows data" report a restart was the only mechanism - and a removed device kept its last good SoC/power feeding the plan (`merge_non_null()` hands back the previous reading on nulls and failed polls) and kept having registers written to. **Post-#5295:** `refresh_devices()` re-reads the device list hourly (`DEVICE_REFRESH_SECONDS`, 60*60, `gecloud.py`; a multiple of 60 because `run()` is called once a minute, and a cycle that finds a change polls and reads settings whatever else is due), `select_poll_devices()`/`apply_devices()` decide what is polled from the fresh list, and adopting a changed set drops every per-device store - `status`, `meter`, `info`, `settings`, `pending_writes`, `register_list`, the EV-charger stores, `ems_slot_warned` and `register_entity_map` - so a removed device stops publishing and being written to. The review round also: the change signature includes `battery_meters`, so a CT rewiring re-runs automatic config; a failed register-list fetch is retried on the next settings read (previously stored as `None` and never fetched again - every later read raised `TypeError`, newly reachable from rediscovery); a device read that would leave zero battery inverters is not adopted (the warning says restart if the hardware really was removed; charger-only sites stay quiet); `ge_cloud_data` returns to its pre-EMS value when the EMS leaves; the EMS slot-override warning re-arms when the EMS changes; and default options for a device discovered while read-only is set stay pending across cycles, applied once read-only ends. Caveats that survive the fix: the 5-day `last_updated` skip inside `async_get_devices()` treats a device silent over 5 days as non-functional wherever that runs - the hourly re-read included - and `merge_non_null()` still hands a *polled* device its last good SoC/power on nulls and failed polls. In-tree fix precedents (GE Cloud itself now among them): Fox re-fetches its device list every 24h (`FOX_REFRESH_STATIC`), SolaX gates plant/device-info re-reads with `data_is_due`, and `num_inverters` is re-read by the core every cycle so a re-run automatic config takes effect next plan. **The "disable" writes on the unused numbered slots leave them live (GH#5371, reporter field log from a GIV-3HY-11 Beta site, code-verified on main):** `enable_default_options()` - run at startup, once per 24h and on newly-found devices, gated only by `read_only_now()` so manual-config GEC sites get the resets too - writes every unused AC-charge and DC-discharge slot 2-10's **times** to `00:00` "to disable" (`gecloud.py`, the numbered-slot branch of the limit/time resets), and on 3-phase GIV-3HY hardware `00:00-00:00` is **live** 00:00:00-00:00:59 (minute bounds), so every "disabled" unused slot is live for one minute past midnight; with force charge on at that minute (a planned charge window spanning midnight, or the v9.1.0 stuck-on case) the unused slots fire at full rate toward their 100% upper limit (discharge mirror: full rate toward the 4% lower floor). The same pass overwrites user-set spare-slot times (the 03:00 workaround) once per pass, per the reporter's inverter-settings history log. Fix direction (maintainer's call): have the **limits** carry the disable instead (unused slots' upper→4%/reserve floor, lower→100%); the constraint is that the two generic limit branches (`*_lower_soc_percent_limit`→4%, `_upper_soc_percent_limit`/`ac_charge_upper_percent_limit`→100%) also match the numbered slots and would write straight back on the next pass, so the fix must special-case slots 2-10 inside those branches - and #5017's outcome could later drive exports through slot 3 on GIV-3HY, so the "unused" set needs a guard there. `test_ge_cloud.py`'s 00:00-reset tests stay valid under this; the fix needs numbered-slot limit assertions added. | `ge_cloud` | | Teslemetry / Powerwall (`teslemetry.py`) | Battery model is inferred from site `nameplate_power / battery_count`. A nameplate fallback combined with `inverter_hybrid: True` once produced a false 5 kW inverter limit and spurious morning export. Tariff writes push a whole TOU schedule every cycle, so a setting changed by hand in the Tesla app is reverted on the next cycle (GH#4600, GH#4610). Three later findings. `build_tariff()`/`_render_side` carve a *single* ON_PEAK interval priced with one scalar (`teslemetry.py:1117`/`:1036`/`:1074`), so no intra-window price gradient can exist and Tesla's own TBC is free to defer the whole export to the end of the window; the no-rate-data fallback branch also skips the boost entirely (GH#4887). The charge path sends no rate at all - `evaluate_schedule` (`teslemetry.py:539`) asserts mode `backup` plus reserve = charge target plus grid charging - so a slow charge ramp is Tesla firmware, not a Predbat write bug (GH#4892); Tesla now snaps a backup reserve of 81-99% down to 80% while Predbat still forwards any 0-100 target. The site_info limit mapping was reworked in PR #5276 (GH#5275, merged 2026-09-28): `inverter_limit` comes from `nameplate_power` (the Powerwall's own AC rating; `max_site_meter_power_ac` is the site's supply limit at the meter, not the inverter's), `export_limit` from `min_site_meter_power_ac` (kW, possibly fractional, with the ±1e9 "no limit" sentinel skipped and 0 published as a real 0 W limit), `inverter_limit_charge` at **5 kW per battery unit** capped at nameplate (Tesla reports no usable charge rating — a Powerwall 3 lists only expansion packs with every rating zero, and fleet data fits the rule), `soc_max` from the gateway's `nameplate_energy_watts` before the `battery_count` estimate, and `inverter_hybrid` now defaults from the Powerwall model (Powerwall 3 on, every other model off) with a `teslemetry_hybrid` override for a PW3 beside an existing string inverter. All of those are wired with `set_arg_auto(overwrite=False)`, so **an apps.yaml value wins** — on a pre-#5276 version `automatic_config()` overwrote a manually set `battery_rate_max` whenever `teslemetry_automatic` was on, which was the standing "my override does nothing" answer on this component. Since PR #4976 there is a Predbat-side lever for exactly that slow ramp: `teslemetry_tbc_control` — **on by default since PR #5188** (GH#5186; `DEFAULT_TBC_CONTROL = True` in `teslemetry.py`, mirrored in `COMPONENT_LIST`, with `teslemetry_tbc_control: False` as the opt-out) — switches `evaluate_schedule()` to `evaluate_schedule_tbc()`, which pushes a **signal tariff** (`build_signal_tariff()` - fixed 0/50/100p bands over the committed charge and export windows) so Tesla's own optimiser runs the charge and reaches the full rate reserve-driven charging cannot; mode goes to `autonomous` in every state, reserve becomes an actual reserve rather than a charge signal, and `_settable_reserve()` maps requests onto reserves the Powerwall honours - an 81-99% request now rounds **up** to 100 rather than being forwarded raw into the snap-to-80 band. The same PR stopped `plan.py` planning a manual freeze export on inverters that cannot do it. **Consequences of the default flip:** a Teslemetry user who never set the key is on the signal tariff with autonomous mode, so "the Tesla app shows 0p/50p/100p instead of my rates" or "my charge target isn't honoured" is the default path working, not a bug; and the 81-99% `set_reserve_min` limitation becomes default-on behaviour with it. (`MockTeslemetryAPI` still defaults `tbc_control = False` on purpose, so a green teslemetry test run exercises the real-rate path unless a test opts in.) And a separate scoping fact from the GH#5186 triage: `teslemetry_automatic` gates only `automatic_config()` — the scheduler emulator (`sync_tariff()` + `assert_device_state(evaluate_schedule(...))`) runs for every non-read-only user with no `automatic` check, so "Predbat changed my Tesla tariff/mode but I never turned on automatic" is that scoping, not a bug; `set_read_only` is the only switch that stops the writes. A related pricing bug in the same `build_tariff()` family was fixed in PR #4977: an export window ending at midnight priced the whole of tomorrow at peak. And `optimization_strategy: "economics"` *is* pushed with `tariff_content_v2` on every `set_tariff()`, unconditionally, since v8.49.0 / PR #4603 - a request to "add" it is asking for something already shipped (GH#4918). **GH#5157** (reporter's evidence, not independently probed — Powerwall 2, fw 26.26.4): the unit sat in the idle branch of `evaluate_schedule` and simply stopped acting on a standing device tuple for ~4h while `site_info` read back the *correct* configuration — drift invisible to every field the API exposes, with only a Predbat restart recovering it; the stall state is the one branch Predbat legitimately holds for hours, which is why transition-based self-heal never fires. A "Powerwall plateaued / stopped discharging until I restarted Predbat" report maps here first. **Implemented on main** (2026-09-19): `run()` re-asserts the whole device tuple — export rule, grid charging, reserve, mode, not the tariff — every `FORCED_ASSERT_SECONDS` (2h) with the write-on-change dedupe bypassed (`force=True`), the timer advancing only on a fully successful forced assert; the Info line `Teslemetry control drift-correction is transition-based ..., backed by a forced re-assert of the full device tuple every 120 minutes` names it, so on current main the stall is bounded at ~2h. **Freeze export is fully implemented in the planner and disabled for TESLA at exactly one point (GH#5225, enhancement):** `INVERTER_DEF["TESLA"]` declares `support_discharge_freeze: False` (`config.py`), read into `inv_support_discharge_freeze` (`inverter.py`), whose only consumer is the first-inverter lockstep block in `execute.py` that force-disables `set_export_freeze`/`set_export_freeze_only` (`Note: Inverter does not support discharge freeze - disabled`) — `EXPORT_MODE_FREEZE` never enters the export ladder, and PR #4976 additionally drops manual freeze-export windows to demand with a Warn. **Flipping the flag alone is not enough:** the component must also map the freeze-target discharge window to a device-side hold, or the plan's freeze assumption fails on the device — the non-TBC `evaluate_schedule()` writes the *raw* discharge target as the reserve, so a 99 target lands in the 81-99 band Tesla silently snaps to 80 (`_settable_reserve()` exists for exactly that and the non-TBC path doesn't route through it), while the default-TBC `evaluate_schedule_tbc()` discharge branch yields `pv_only` with the standing reserve but no hold — the battery still serves house load and drifts down. The device primitive exists in-tree: the TBC charge-hold tuple (`pv_only` + reserve 100 + grid charging off + autonomous) holds the battery while PV surplus exports, but a freeze *export* hold must sit at the current SoC (reserve at/above SoC through `_settable_reserve()`), not 100 — reserve 100 lets solar recharge the battery instead of exporting the surplus. Every `INVERTER_DEF` entry with `support_discharge_freeze: False` (the cloud-emulator types: SunsynkCloud, AlphaESSCloud, DeyeCloud, SolaxCloud, SolisCloud…) shares this two-part requirement — capability flag **plus** device-side freeze mapping — so enabling freeze export on any of them is a two-file change, not a one-line flag flip; each component's schedule emulator needs its own hold tuple designed against what its hardware honours (*suspected*, not verified per component). | `teslemetry` | @@ -117,7 +117,7 @@ Grep for the named symbol rather than trusting a line number. | Grid sign / arrow direction | `grid_power_invert` is owned by some integrations and not others. With two systems configured, one integration setting it `True` bleeds into the other's entities and inverts the arrows. The fix is an explicit `False` in the automatic config of both. | `sunsynk_config`, `teslemetry` | | Enphase (`enphase.py`) | Unofficial Enlighten endpoints. Accounts with MFA cannot log in at all. Discharge-to-grid schedules are required for export control. Writes need a double-submit CSRF token or return 403. Using the Enphase app at the same time can trip session limits. | `enphase_api` | | myenergi (`myenergi.py`) | `myenergi_automatic_zappi` (PR #4997) gates the Zappi half of automatic config; with it off, Zappis stay monitor-only and `control_active` is never set. `release_zappis()` has exactly one caller — `control_tick()` — reached only under `if self.control_active:` in `run()`, and `control_active` is latched at the first tick: every `enable_control()` refusal (zappi_control → automatic → automatic_zappi → enable_controls) returns *before* setting it, and the gate re-runs only when a fresh instance starts. So after a component restart with any prerequisite off, a Zappi already held in 'Stopped' stays Stopped and the `switch.*_myenergi_zappi_control` entity is not even republished (it is gated on `control_active` too) — the symptom for a future report is "car won't charge after I turned off Predbat" with a Zappi. Note `control_tick()` itself *does* release on read-only mode and on the control switch being turned off; it is the config-gate half that strands. Since PR #5150 (merged 2026-09-19) sunsynk/deye instead persist `control_active` across a restart (and re-infer it from a pre-upgrade cache), so a restart cannot strand an already-armed inverter there — the myenergi strand is the opposite direction (the gate never latching in the first place) and remains. Two adjacent lifecycle facts: published entities are **never removed** — `dashboard_item()`/`set_state()` are upsert-only (`output.py`, `ha.py` POST `/api/states`), so when a gate stops `publish_data()` publishing a switch the old entity lingers frozen at its last state until HA restarts, and there is no entity-removal path anywhere in the codebase; and `set_arg`/`set_arg_auto` writes survive a component-only restart, because `Components.restart()` stops and re-creates only the instance and never resets the shared `base.args` from apps.yaml — auto-wired values (e.g. myenergi's `car_charging_*`) persist until a full process restart. Distinguish the two restarts when checking arg lifecycle. | `myenergi` | -| Gateway MQTT (`gateway.py`) | Control writes are MQTT commands the hub acknowledges. **The ack design arrived with PR #5220 and was refined by PR #5231 (both merged 2026-09-26)** — both postdate every earlier gateway log, so keep a publish-per-call reading for pre-#5220 logs: `_subscribe_acks()` subscribes to `predbat/devices//ack/+`, tolerating a broker that refuses it (a single warn via `_ack_subscribe_warned`; without the subscription or before any ack, `_send_control()` publishes exactly as before); once acks are seen, an identical command (same entity, command, payload) is published once per `_COMMAND_ACK_WINDOW` (30s), re-sent once after the unanswered window, and if that re-send is also unanswered `_acks_seen` drops so writes fall back to publish-every-call until telemetry confirms; `_process_ack()` matches any tracked id of the current value, a refusal outranks an ok (a multi-unit dispatch acks once per unit), and replay errors are handled; ids come from a clock-seeded monotonic counter (`_next_command_id()`, `PBAT`), so they never restart at `PBAT1` after a restart, and the kept-id list (`_COMMAND_ACK_IDS_KEPT`) is pruned only after the new send is out, with an ack that arrives during a failing publish still counting. **PR #5172 is a separate, still-open implementation of the same feature** (`_subscribe_command_acks`, `uuid4().hex` ids, typed outcomes threaded through service dispatch/HTTP/inverter verifier) — its symbols are not on main until it rebases, so do not start a gateway-ack triage from its names. **Two gaps around serial-less and empty status, probed on main 2026-09-25 (GH#5227, enhancement — probed with temporary tests in `test_gateway.py`, reverted after):** (1) a sole slot publishing `serial=""` (shipped firmware 1.0.0 publishes every slot, including a GivEnergy slot whose serial discovery failed) is bound as the control target: `_needs_reconfigure()` treats `""` as a newly discovered inverter, the last-resort branch (`candidate_aios or list(all_inverters)`, `gateway.py`) binds it, and the result is empty-suffix entities (`select.predbat_gateway__charge_slot1_start`) and commands addressed to `""`, which the firmware rejects (`dongle_serial required`) — a real fleet incident on 2026-09-16 (~22 min of failed commands); an empty-serial guard in `_needs_reconfigure()`/`automatic_config()` would close it, and #5227's proposed `dongle_count` deferral does *not* cover this state (`dongle_count == len(inverters)` when the empty slot is the only one). (2) `if len(status.inverters) == 0: return` (`gateway.py`) fires **before** `_inject_entities()` (EV chargers, gateway-online) and before `_last_telemetry_time`/`update_success_timestamp()` — on a hub where all GivEnergy slots are withheld (post-predbat-gateway#335, e.g. a single-inverter hub), every status message is dropped: EV data freezes and the component fails the 60-minute staleness test (`components.py`). Symptom pointer: "gateway EV sensors froze / gateway component unhealthy while the hub is up" → this early return. Removals staying bound and reappearance triggering re-config are already the behaviour (probed). **Probe trap:** the topology branch in `automatic_config()` depends on the *whole* visible set — a probe that adds a Gateway alongside the `""` slot drops the `""` entry via the battery-presence filter (`aios` requires `battery.ByteSize() > 0`) so it never reaches the last-resort branch; construct the exact visible set the scenario implies, and note the filter silently hides data-less slots from the count in several branches (the same mechanism can make a Gateway+2-AIO fleet collapse to AIO-direct control). Existing precedent for topology probes: `TestGatewayUnitControlBinding` (`test_gateway.py`). | `gateway` | +| Gateway MQTT (`gateway.py`) | Control writes are MQTT commands the hub acknowledges. **The ack design arrived with PR #5220 and was refined by PR #5231 (both merged 2026-09-26)** — both postdate every earlier gateway log, so keep a publish-per-call reading for pre-#5220 logs: `_subscribe_acks()` subscribes to `predbat/devices//ack/+`, tolerating a broker that refuses it (a single warn via `_ack_subscribe_warned`; without the subscription or before any ack, `_send_control()` publishes exactly as before); once acks are seen, an identical command (same entity, command, payload) is published once per `_COMMAND_ACK_WINDOW` (30s), re-sent once after the unanswered window, and if that re-send is also unanswered `_acks_seen` drops so writes fall back to publish-every-call until telemetry confirms; `_process_ack()` matches any tracked id of the current value, a refusal outranks an ok (a multi-unit dispatch acks once per unit), and replay errors are handled; ids come from a clock-seeded monotonic counter (`_next_command_id()`, `PBAT`), so they never restart at `PBAT1` after a restart, and the kept-id list (`_COMMAND_ACK_IDS_KEPT`) is pruned only after the new send is out, with an ack that arrives during a failing publish still counting. **PR #5172 is a separate, still-open implementation of the same feature** (`_subscribe_command_acks`, `uuid4().hex` ids, typed outcomes threaded through service dispatch/HTTP/inverter verifier) — its symbols are not on main until it rebases, so do not start a gateway-ack triage from its names. **Two gaps around serial-less and empty status, probed on main 2026-09-25 (GH#5227, enhancement — probed with temporary tests in `test_gateway.py`, reverted after):** (1) a sole slot publishing `serial=""` (shipped firmware 1.0.0 publishes every slot, including a GivEnergy slot whose serial discovery failed) is bound as the control target: `_needs_reconfigure()` treats `""` as a newly discovered inverter, `_needs_reconfigure()` treats `""` as a newly discovered inverter and auto-config binds it — note the last-resort branch that used to fall back to `candidate_aios or list(all_inverters)` was removed in `eaef6e61` (2026-10-04), which now defers auto-config entirely with a warn when no battery-capable telemetry has arrived (`_auto_configured` stays False so `_needs_reconfigure()` retries); an empty-serial slot that does report battery telemetry is still bound, so the trap survives on that shape, and the result is empty-suffix entities (`select.predbat_gateway__charge_slot1_start`) and commands addressed to `""`, which the firmware rejects (`dongle_serial required`) — a real fleet incident on 2026-09-16 (~22 min of failed commands); an empty-serial guard in `_needs_reconfigure()`/`automatic_config()` would close it, and #5227's proposed `dongle_count` deferral does *not* cover this state (`dongle_count == len(inverters)` when the empty slot is the only one). (2) `if len(status.inverters) == 0: return` (`gateway.py`) fires **before** `_inject_entities()` (EV chargers, gateway-online) and before `_last_telemetry_time`/`update_success_timestamp()` — on a hub where all GivEnergy slots are withheld (post-predbat-gateway#335, e.g. a single-inverter hub), every status message is dropped: EV data freezes and the component fails the 60-minute staleness test (`components.py`). Symptom pointer: "gateway EV sensors froze / gateway component unhealthy while the hub is up" → this early return. Removals staying bound and reappearance triggering re-config are already the behaviour (probed). **Probe trap:** the topology branch in `automatic_config()` depends on the *whole* visible set — a probe that adds a Gateway alongside the `""` slot drops the `""` entry via the battery-presence filter (`aios` requires `battery.ByteSize() > 0`) so it never reaches the last-resort branch; construct the exact visible set the scenario implies, and note the filter silently hides data-less slots from the count in several branches (the same mechanism can make a Gateway+2-AIO fleet collapse to AIO-direct control). Existing precedent for topology probes: `TestGatewayUnitControlBinding` (`test_gateway.py`). | `gateway` | | Octopus (`octopus.py`, `fetch.py`) | Intelligent Go tariffs are detected via `is_intelligent_go_tariff()`, and IOG-prefixed tariffs must be skipped when updating intelligent devices. Saving-session auto-join rebinding regressed when `joined_events` was empty (GH#4573). `octopus_slots_signature()` deliberately omits the time-drifting fields of active dispatch slots so a replan is not forced every cycle. `car_charging_threshold` is a strict fallback gated on `not self.car_charging_energy` (`fetch.py:228`, and `load_ml_component.py:348-361`) — it never runs as a second filter alongside a real `car_charging_energy` sensor (GH#4717). Saving-session reporting credits the full saving rate to every minute of the session on both rate tables (`load_saving_slot()`, `octopus.py:2786`); the slot dict has no baseline field to subtract (GH#2090) — a complaint that a saving session's reported total looks inflated starts here, not in a rate-fetch bug. A tariff with no `standard_unit_rates` link (IOG-TOU, GO) goes down `async_get_day_night_rates()`, which infers the off-peak window from 7 days of measurement TOU labels by plurality vote. On IOG those labels include ad-hoc dispatch slots, so a bonus slot recurring on 4+ of 7 nights became a permanent nightly cheap window and Predbat charged into it at the day rate (PR #4854 skips the inference for IOG). A "cheap slot Octopus has never heard of" report where the dispatch feeds are *empty* is this, not phantom dispatches — check the log for `Using off-peak windows [...] from measurement TOU labels`. TOU labels are UTC instants, so windows derived from them must be re-anchored for a local wall-clock tariff or they drift an hour at DST. Free-session slots have their own family of gates, separate from the saving-session ones and each found the hard way. `octopus_free_session` events used to be dropped when `code` was null, which is exactly how Octopus publishes auto-joined Weekend Happy Hours - fixed, the gate is now `start and end` and the log falls back to the event id (GH#4835). The free/saving-session args name **event** entities (`attribute="events"`/`"joined_events"`, `octopus.py`), and the BottlecapDave integration exposes Octoplus sessions as calendar + event pairs of which only the event entity carries those state attributes - the calendars have none at all, and both sides ship **disabled by default** in the integration. "The documented sensor name does not exist" plus a proposal to point at the `calendar.` twin is almost always the disabled-default trap, not a rename (GH#5370, verified against the integration's source): enable the event entity; the calendar twin cannot back these args. Matching calendar/event names differ by the `_events` suffix, so a pattern built from one domain's names mistranslates, and the deprecated pre-v17 names are scheduled for removal in January 2027. `load_free_slot()` bounded itself with `start_minutes < self.forecast_minutes` while rate minutes are indexed from `midnight_utc`, so any event more than `forecast_minutes` past midnight was dropped with no log line at all - fixed in `3bacc6e8`, the gate and the end clamp now both use `forecast_minutes + minutes_now` like `load_saving_slot()` (GH#4931). In the same function the start/end range was only updated on a successful decode while the apply block ran unconditionally, so an undecodable slot re-applied its rate over the *previous* slot's minute range (PR #4935). The join side has its own: the `joined_events` guard still ends in `saving_rate > 0` (`octopus.py:3542`), so `octopoints_per_kwh: 0` - a genuine free hour - yields no slot of any kind; probe-verified (0 gives nothing, 500 gives a saving slot, null gives nothing), and `git blame` puts that guard in `1b8c9136`, the fix for null-rate sessions (GH#3079), so it is an oversight rather than a deliberate exclusion (GH#4851). On dispatches, `rate_add_io_slots()`'s `location` check used to apply to *completed* dispatches too, and `rate_import` is rebuilt every cycle, so a retroactive AT_HOME to AWAY relabel silently un-stamped the cheap rate and `today_cost()` re-priced the whole day at the day rate (GH#4946) - **fixed in PR #4957**, which added `dispatch_billed_off_peak()` (`octopus.py:2977`): a completed dispatch is billed off-peak regardless of location, a straddling one keeps the location test, and the midday budget cap still applies. Keep the mechanism in mind when reading a log from before that merge, where the whole day reprices at the day rate. A neighbouring cap bug had no entry here at all: the daily low-rate slot budget was keyed on the *loop* minute rather than the slot start, so a window straddling noon drew a fresh 12-slot budget at 12:00 and stamped the afternoon cheap (GH#4950, **fixed in PR #4951**). The midday boundary itself is deliberate (`337db867`); only the keying was wrong. Worth knowing because "a cheap slot at an hour Octopus never offered" has at least three distinct causes in this row alone. Compare has its own rate-fetch gap: `download_octopus_rates_func()` (`octopus.py:2725`) reads only `standard-unit-rates`, and the newer IOG-SMB-FIX products are `four_rate_ev` - that endpoint answers 200 with empty results and the rates live on `day-unit-rates`/`night-unit-rates`. The main OctopusAPI component already auto-detects that product shape; compare never got the same treatment (GH#4921, probed against the live API). Lastly, `{dno_region}` in a tariff URL is substituted by `resolve_arg`'s `.format(**self.args)` (`userinterface.py:106-116`), so a missing `-` before the placeholder glues the region letter onto the product code and yields a plausible-looking 404 - check the literal URL in apps.yaml before believing a tariff has been withdrawn. Two later additions to the free/saving-session picture, one of them a correction. **eventType is carried only by the legacy `savingSessions` query path — the flexibility feed drops it (GH#4548, re-verified on main 2026-09-29 by reading `async_get_flexibility_events()`):** that function extracts only `code`/`startAt`/`endAt` from `customerFlexibilityCampaignEvents` and maps every saving event into `joinedEvents` with **no eventType**, so on the Direct path with an MPAN the joined-event loop's free-slot branch (`event_type == "WEEKEND_HAPPY_HOUR"`) can never fire and a joined Happy Hour falls through to the rewarded-saving-slot branch whenever `octopus_saving_session_rate` is set above 0. The legacy query carries eventType (the `savingSessions` GraphQL request names it and its map stores it) and is reached on the Direct path only as the flexibility path's own fallback (no MPAN, or the flexibility API returned no saving events). eventType remains the only sound discriminator between a Power Down saving session and a free Power Up/Happy Hour — the field GH#4851's `saving_rate > 0` problem needed and did not have — but *which query path served the events* decides whether you have it: "free Power Up priced as a paid saving session" on the flexibility path is first a question of provenance, and waiting for the BottlecapDave API before building on the newer feed is deliberate, not an oversight. The Direct join mutation is likewise still the legacy `joinSavingSessionsEvent`. And a genuinely counter-intuitive ordering constraint, worth reading before touching that function: Weekend Happy Hours are now skipped from `available_events` (they cannot be joined through the API - Octopus allocates them or the user books on the website), but **the skip has to sit after the reward/code/type maps are populated**, because those maps are built from the same events list and a joined Happy Hour looks its own type up there. Skip too early and the joined event has no type, takes the injected default reward, and the planner prices a free hour as an **80p/kWh saving session** (`c0fb4e9c`; the test fails if the skip is moved above the maps). Two rate-provenance reports from mid-September. **GH#5012** (car unplugged, plan still charged the house battery in the phantom cheap window): `octopus_intelligent_ignore_unplugged` correctly removed the planned dispatches, but the cheap price came from the integration's own rate sensor, not Predbat's stamping — parse the debug yaml and compare, per suspect minute, `rate_import_base` vs `rate_import_no_io` vs `io_adjusted`: a cheap price present in base and no_io with `io_adjusted` empty is the integration's rate sensor; present only after no_io is Predbat's `rate_add_io_slots()` stamping (#4516/#4950 family); `io_adjusted` populated means the minutes were flagged as adjusted — but **since PR #5304 (v9.3.3) Predbat's own `rate_add_io_slots()` flags the minutes it lowers too** (it used to write no marker at all; it mirrors the integration's `is_intelligent_adjusted`, never the fixed off-peak, time over the daily cap, or a cancelled car), so a populated `io_adjusted` no longer discriminates the integration's stamp from Predbat's own; the base-vs-`no_io` comparison still does, and `rate_replicate()` still refuses to copy `io_adjusted`-flagged minutes into future days. This attributes "integration feeds the wrong price" vs "Predbat stamps the wrong price" in one step. Nothing reverts `io_adjusted`-flagged minutes when the car is unplugged + ignore_unplugged is on — that fix would be an enhancement, blocked in practice by the integration not flagging stale adjusted entries (`io_adjusted` was empty in this dump). **`io_adjusted` itself was for a while wiped by the export/gas fetch (GH#5286, fixed in PR #5290, merged 2026-09-28):** `fetch_octopus_rates()` replaced `self.io_adjusted` on every call, so the export and gas fetches that follow the import fetch reset it to `{}` whenever `metric_octopus_export`/`metric_octopus_gas` were set — regression from #2826 dropping the old "only when adjust_key is set" guard — and with no markers left, `dynamic_load_car_strip_feed_rates()` had nothing to strip, so a car cancelled out of a dispatch still left it priced cheap and the house battery planned into it. The fix makes `self.io_adjusted` replaced only when the fetch carries an `adjust_key`; keep the mechanism for pre-#5290 logs. A companion fix (same PR) stops a compared tariff inheriting the live tariff's dispatch markers — each compared tariff now starts with none, and one left on the live import rates gets a copy. A user who instead wants the house battery limited to minutes the car actually charges has no clean knob (GH#5065): `octopus_intelligent_charging` off does **not** stop the stamping — the switch gates only car planning and vehicle prefs, planned slots are collected into `self.octopus_slots` regardless (gated only by `octopus_intelligent_ignore_unplugged`, which covers unplugged, not plugged-in-but-deferred) and are consumed unconditionally by `rate_add_io_slots()`. `octopus_slot_max: 0` is not selective either — the cap is applied in `load_octopus_slots()` as well as `rate_add_io_slots()`, so it removes the car's charging slots too. The only complete workaround is unsetting `octopus_intelligent_slot`, which loses completed-slot tracking and Octopus car planning with it. **GH#5018** (tariff switched mid-day; tomorrow's export rates stayed on the old tariff and Nordpool estimates never appeared): Predbat re-reads Octopus rate events from HA every fetch cycle (`fetch_octopus_rates`), so stale export rates are the HA integration still serving the old contract — reload the integration. Two Predbat-side traps in the same report: `futurerate` only fills *missing* minutes (the `if minute not in rates` gate in fetch.py), so stale-but-present rates always win over Nordpool estimates; and `futurerate_adjust_auto` re-runs at every FutureRate init (one per fetch cycle) and persists via `set_arg`, so mid-switch it can persist `False` over the user's manual `futurerate_adjust_export: true` — log signature `FutureRate: No futurerate adjustment enabled, skipping futurerate analysis` repeating every cycle. The debug yaml does **not** contain the Octopus rate events (grep for `day_rates` comes up empty), so replays can't reproduce integration-served rate data; use the log's `Export rates: min/max/average` lines as the provenance check (identical min/max across days = static served data, and a fixed tariff has a fixed shape while a variable one doesn't). `self.mpan` is the **import** MPAN only — `async_find_tariffs()` sets it once from the first active import agreement's `meterPoint` and never overwrites it; an account with export has import and export as separate agreements with different MPANs, and the export one is retained nowhere (PR #4972 review). The saving-session/free-electricity GraphQL queries use `self.mpan` deliberately because those campaigns are import-side; any feature that needs to identify a *meter* cannot take `self.mpan` (PR #4972, merged 2026-09-19, adds `self.tariffs[direction]["mpan"]`). Symptom: two things that should describe different supply points coming out identical — that is this, not a deduplication bug. **GH#5144** (fixed in PR #5145): the REST tariff endpoints return overlapping `DIRECT_DEBIT`/`NON_DIRECT_DEBIT` rows for the same validity window and `minute_data()` writes each row over its range, so whichever row came last in the response won — and the order is not stable across periods, so the displayed rate could flip between variants from one period to the next ("always the higher rate" was an artefact of that ordering, not a property of the bug). `filter_payment_method()` (`utils.py`) now keeps one variant at the parse point on the component, day/night, annual and minute-data paths — preferring `DIRECT_DEBIT`, keeping rows with no `payment_method` (Agile) untouched, and leaving single-variant tariffs unchanged; null `payment_method` is the common case, so any future filter must keep nulls. **The daily cheap-slot cap is derived per account but enforced per car (GH#5215, semantics unresolved):** `get_octopus_slot_max()` (`octopus.py`) takes no `car_n` — it resolves the cap from the *account's* import tariff code via `has_six_hour_cap()` (matches `IOG-SMB` → 12) or a single apps.yaml integer — while both enforcement counters are locals (`slots_per_day` in `rate_add_io_slots()` and in `load_octopus_slots()`), and both callers run once per car, so each car draws a fresh budget: up to `octopus_slot_max × num_cars` cheap half-hours priced per day. Do **not** assume the docs' per-car sentence is simply wrong: the derivation is account-level but Octopus's own blog says "*Your car gets up to 6 hours of off-peak charging per day*" (quoted in GH#4830, the multi-car confusion thread), so the code is internally inconsistent and the intended semantics is genuinely open. Exposure is invisible outside IOG-SMB — uncapped tariffs default `octopus_slot_max` to 48 — so only capped multi-car installs can see the doubling; existing tests encode the per-car behaviour, and `octopus_slot_max` is read without `index=` so per-car configuration does not exist either way. **The IOG slot-confirmation strip removes the house's cheap rate too when a car's dispatch is cancelled (GH#5335, with the #5317 family - code chain on main 57ec7bf1):** with `octopus_intelligent_dynamic` on (default), `dynamic_load_car_check()` (`plan.py`) cancels every slot of a car that sits inside a started dispatch but is not seen charging after the confirmation grace (log: `car 0 is in a dispatch but not charging, cancelling its slots`), and `dynamic_load_car_strip_feed_rates()` (`octopus.py`) re-prices the house's view of those minutes back to `rate_max_base` when the cheapness came from the integration feed, **exempting the fixed IOG band** (`OCTOPUS_NIGHT_RATE_WINDOWS["iog"]` = 23:30–05:30) — so a house battery that had planned into the dispatch loses its cheap stamp as well, and PV10 re-prices the minutes at `rate_max` outright (prediction.py's `dispatch_gone` line, `minute > 30`). Version tell: the strip is v9.3.2+ code — a reporter's deliberate mid-log downgrade left the A/B in one `predbat.log`: v9.3.3 logged `Octopus Intelligent: removed the dispatch rate from 355 minutes of cars [0] which are not charging` while v9.3.1 kept the 6.57p dispatch rate (`Charging target 2%-100%`); workaround (semantics read, not live-tested): `octopus_intelligent_dynamic` off = no confirmation and no strip, ≈9.3.1 behaviour at the cost of #5229's low-load protection. Check the car side first, though: Predbat must *see* the car charging to confirm the slot, and a `plug_status` string (`eco`/`boost`) that never equals `car_charging_now_response: charging` leaves the car "not charging" through a real charge — check the sensor's own history before blaming the planner. **#5316 is the mid-dispatch variant, fixed in PR #5319 (merged 2026-10-03) — keep the pre-#5319 signature:** a car that stops part-way through a running dispatch half hour used to have the rest of that half hour re-priced at the day rate; the strip paths (`rate_add_io_slots()`, `dynamic_load_car_strip_feed_rates()`) now start from `dynamic_load_car_strip_from()` (`plan.py`) — the end of the half hour the car was last seen charging in, which is billed off-peak in full by Octopus — and the confirmation is never cleared, only outlived. Review-round refinements ride along: a sensor reading confirms a half hour only from 2 minutes into it, and the grace clock is not kept across a restart (saved state is judged on the skewed clock). | `octopus_*`, `saving_session*` | | Kraken / EDF (`kraken.py`) | EDF Kraken answers a day/night-structured tariff with **HTTP 400** on `standard-unit-rates/` (`{"detail": "This tariff has day and night rates, not standard."}`) while `day-unit-rates`/`night-unit-rates` return 200 and standing charges 200 — where Octopus four-rate products (GH#4921) answer the same-shaped URL **200 with empty results**. So the "empty results, then data empty" signature belongs to `octopus.py` and a hard 400 to `kraken.py`; do not carry the expectation across (GH#5166, probed against the live API). A nonexistent product still 404s, so 404 and 400 are both live "REST can't serve this tariff" signals on EDF — probe the exact URL from the reporter's log before assuming either. **Fixed in PR #5167** (`async_fetch_rates()`): the GraphQL `applicableRates` fallback now gates on `KRAKEN_REST_RATES_UNAVAILABLE_STATUSES` = (400, 404, 410), per direction (404 covers both a private product and one retired from the REST API; the authenticated retry stays 404/410 because auth cannot change a 400), and only counts a failure when the fallback comes back empty without counting one — `async_graphql_query()` counts its own failures, so a caller that also counts scores 2 per down cycle; snapshot the counter around the call. The standing-charge fallback stays deliberately 404/410-only (`kraken.py` — "400 is a rates-endpoint-only answer"). Two traps from the fix's review: `_fetch_rates_rest()` returns `(None, None)` on a network error, so `err is None` does **not** mean success — success is `(results, None)`; the merged code stashes the public attempt's status and restores it when the authenticated retry hits a network error (pre-fix symptom: a private-product tariff's rates vanish for one cycle with no fallback log line, while an identical 404 one cycle later recovers). The day/night endpoints' rows carry the same `payment_method` variant overlap as GH#5144 — suspected, no probe; check `minute_data`'s handling before consuming them directly. Also benign: `Warn: Kraken: Auth not available for find-tariffs` repeating on early restarts before the first `Tariff discovered` line is setup-phase, not a second bug. | `kraken` | | Axle (`axle.py`) | Export sessions have to boost the import rate as well as the export rate - introduced deliberately by PR #4520 (first release v8.48.2, mirrored on Octopus saving sessions), documented in docs/energy-rates.md, and upheld by the maintainer on #5060. Two reports of the same design now exist (#5060 closed as dup-of-#4277 with the by-design answer given; #5175 triaged as enhancement, no regression - `tools/triage_test.sh axle` asserts the dual boost), and **no config knob disables just the import side** (`axle_pence_per_kwh` scales both directions; config.py has only `axle_api_key`/`axle_pence_per_kwh`/`axle_automatic`/`axle_control`). The "textual plans / colour coding skewed" half of such reports belongs to the #5050 threshold row (boosts before `rate_scan()`), not here. State is published unconditionally from `run()` so a fetch failure does not freeze the sensor at a stale value. `load_axle_slot()` has no lower time bound - it checks only `start_minutes < forecast_minutes + minutes_now`, not the `start_minutes >= 0` guard `load_free_slot()` has (octopus.py) - so a closed Axle event writes dead keys at negative minutes through `rate_dict.get(minute, 0)` (GH#5036, replay-verified against the reporter's dump: the +100 boosts sit only at the two closed events' actual UTC start/end times). The dead keys are inert in the plan - `rate_minmax()` and the window scans start at `minutes_now` - so "a dead event projected forward as free export" is not what the code does. The comment at octopus.py:2859 saying `load_saving_slot()` and `load_axle_slot()` "both bound themselves that way" is half-wrong: axle bounds only its end, and `load_saving_slot()` also lacks the start guard but is harmless there because its inner loop writes only `if minute in rate_dict`. Two September reports extend this row. **GH#5060** (closed as a duplicate of #4277, the same defect reported a year earlier and never root-caused — #4277 covers Octopus saving sessions too): event boosts are applied *before* user rate overrides on both sides — `load_axle_slot(..., export=True)` runs ahead of `basic_rates(rates_export_override, ...)` (import mirrors it: `load_saving_slot`/`load_axle_slot` then the import override) — so a full-horizon `rates_export_override` rewrites the event minutes and re-marks them `user`. That reporter had no `rates_export` key in apps.yaml at all, so the override was their only export source, which is why the import-side boost survived while the export-side +100 did not. Workaround (maintainer's suggestion on #4277, checked viable against the code): move the fixed schedule from `rates_export_override` into `rates_export` — it becomes the base tariff via `basic_rates` before the boosts run, so the event stacks on top and nothing downstream rewrites it (assuming manual export rates are empty, since `apply_manual_rates` also runs after the boosts). Forensic technique: in a debug yaml, the `rate_*_replicated` dict discriminates the two failure modes in one step — mark `user` through the event window ⇒ the boost ran and a later override clobbered it; the `saving` marks missing while the boost is missing ⇒ the boost call never ran (or, per GH#5036 above, wrote dead keys). Compare against `rate_*_base` for the tariff's own price; same idea as the GH#5012 base/no_io/adjusted comparison in the Octopus row. **GH#5050** (shared with the Octopus row): the same boosts run before `rate_scan()`, so the automatic low-rate threshold classifies the whole day low — see the symptom-table row on `set_rate_thresholds()`. | `axle` | @@ -130,9 +130,9 @@ Grep for the named symbol rather than trusting a line number. | Predheat (`predheat.py`) | GH#4670: with `predheat_enable` set, Predheat still did not activate after startup because of lazy flag initialisation. There is no registered Predheat test module, so there is nothing to run here — investigate by reading. GH#4848 mapped the surface for the recurring "make Predheat drive my heat pump" requests: Predheat has **no control path at all** - its only outbound service call is a read-only `weather/get_forecasts`, everything else is `set_state` publishing prediction sensors. `smart_thermostat` looks like control but only pre-empts an already-scheduled setpoint rise *inside the simulation*. And Predheat is instantiated as its own object with its own timer loop rather than as a `PredBat` mixin, so `plan.py` never sees a heat decision variable. `load_forecast: - predheat.heat_energy$external` does surface heat consumption as the plan's Xload column, but as a *fixed* load the battery plans around, never a shifted one. **GH#5153** (code + repro verified): `get_weather_data()` passes the forecast's `temperature` raw into `minute_data()` with no unit conversion — a °F-native weather entity feeds ~55-66 into a °C model, the loss term `heat_loss_watts * (internal - external)` then *heats* the house toward the modelled outside temperature and the internal forecast runs away and levels just below it while energy/cost stay 0 because the thermostat correctly never fires. One-glance check: the weather entity's `temperature_unit` attribute. Same issue, second bug: `minute_data_age` includes the age of `heating_energy`, which is optional — unconfigured, its age is 0, `minute_data_age` pins to 0 and `get_historical()` returns its 20.0 default for the target sensor, so the live target is ignored whenever heating_energy is unset (the template includes it; ASHP users without a heat meter remove it). Traps: `run_simulation(save="best")` publishes HA entities by default, so a parameter sweep overwrites the published sensors unless `save` is set otherwise; predheat reads its own clock (`datetime.now(local_tz)`, not the fixture's pinned `now_utc`), so history mocks must be built against the real wall clock or `minute_data_age` comes out negative; and the registered `test_predheat` stubs `update_pred`, so no simulation physics is under any test. | none | | Load ML (`load_ml_component.py`, `load_predictor.py`) | The 2-hourly retrain isn't one pass. `_do_training()` hardcodes `ml_curriculum_step_days = 1` / `ml_curriculum_max_passes = 4` (`load_ml_component.py:94-95` — plain instance attributes, no `config.py` entry, not user-configurable), and `train_curriculum()` caps to the largest N windows (`load_predictor.py:1560-1564`); on ~80 days of history that's roughly five near-full-scale passes back to back, confirmed against a reporter's log as a 16-minute CPU spike every 2 hours (GH#3896). Separately, `threads` only ever reaches the C++ prediction kernel (`plan.py:1497`, `resolve_batch_threads`) — never wired to ML — and there is no `OMP_NUM_THREADS`/`threadpoolctl` anywhere in the repo (confirmed absent, 2026-08-27), so NumPy's BLAS backend is free to fan out across every core on its own during that training window. Whether that actually saturates a host depends on which BLAS backend is linked — worth confirming the platform before assuming a code fix is the right lever. There is a mapped test now: `ml_training_perf` (PR #4912), which asserts the training invariant rather than peak RSS — a peak-memory assertion was too machine-dependent to hold. Frequency is a separate axis (GH#5072): `RETRAIN_INTERVAL_SECONDS` (2h, `load_ml_component.py`) is the only knob of the cadence and is hardcoded, and any "make the retrain interval configurable" implementation hits a silent ceiling at the equally hardcoded `ml_max_model_age_hours = 48` — `LoadPredictor.is_valid()` returns invalid with reason `"stale"` past it and ML predictions then fall back to empty, so exposing the interval without also exposing the staleness cap is a footgun above 48h. Model-status signals (GH#5075, code-verified on main; PR #5112 open draft carries the fix): `is_valid()` can report **"active" for a model that has never been trained** — `_initialize_weights()` sets `model_initialized = True` before the first epoch, and both `validation_mae` and the age check skip `None`, so a training abandoned at epoch 0 publishes `model_status = "active"` / `model_valid = True`; conversely `training_timestamp is None` is *not* a sound "never trained" test, because a legacy saved model that predates the field is trained and must stay valid (`test_load_ml.py` asserts exactly that). And `train()` stamps `training_timestamp`/`validation_mae` at the end of *every* curriculum pass, so a curriculum that aborts part-way leaves the last intermediate window's **fresh** stamps standing — suppressing exactly the `ml_max_model_age_hours` staleness retrain that would otherwise rebuild a complete model. #5112 adds a `model_trained` flag plus stamp snapshot/restore around `train_curriculum()`; until it merges, treat "status active + nonsense forecast on a just-started install, or whose initial training keeps failing" as this, not as a data problem. The normalisation-statistics pairing on an abort (restored weights, stats refit for the abandoned run) is *suspected, not measured*. **"Reverting fixed the CPU" is usually an observation-window artefact (GH#5213, verified from a Pi 5 reporter's log):** load-ML fine-tune ran on *both* versions with identical structure (5 passes, 30 epochs, ~14s/epoch — 20–41 min every 2h ≈ 19% duty of one core), while both of the reporter's short v9.0.3 windows contained a training their next upgrade killed partway, so the panel always read low when they checked after rolling back. On any "CPU regressed after upgrade to X" report with Load ML on, first reconstruct from the log when fine-tunes started/finished under each version and how long the user actually stayed on each — the 2h retrain boundary means a version window shorter than that proves nothing. **Load ML has a hardcoded 48h horizon and the plan's load is zero beyond it (GH#5336, enhancement, read on main 57ec7bf1 2026-10-01):** `PREDICT_HORIZON = 48 * (60 // CHUNK_MINUTES)` (`load_predictor.py:35`) is the only horizon definition, and it feeds only the `predict()` rollout loop, blend schedule, logs and the `save()` metadata — training does not depend on it, and `load()` validates only model version + architecture, so models survive a horizon change. With `load_ml_source` set, `fetch.py` sets `load_forecast_only = True` and the weighted-bucket historical forecast is skipped wholesale, so the plan's future load is `self.load_forecast` alone (via `step_data_history(..., load_forecast=...)` from `fetch_sensor_data()`'s call path) — past minute 2880 there is no data and the plan-HTML load column is exactly 0; the log tell is `Starting autoregressive prediction loop for 576 steps (48.0 hours)`. The "over 48h" in `Generated N predictions (total X kWh over 48h)` (`load_ml_component.py:695`) is hardcoded in the *log string* — the step count is the truth — and `test_load_ml.py` pins 576 steps. The ask (configurable 48/72/96) is architecturally straightforward because little else depends on the constant, and there is no fallback beyond the ML horizon by design of `load_forecast_only`. | `ml_training_perf` | | Savings & metrics (`output.py`, `predbat.py`) | `savings_total_predbat` accumulates the **unadjusted** `saving` (`predbat.py:1183`, fed from `self.savings_today_predbat = saving` in `output.py`) while `savings_yesterday_predbat` publishes `saving_adjusted` (`output.py:3380`) — the two sensors answer different questions and can disagree in sign on the same day (confirmed from a reporter's dump: `saving_real: +86.14p` vs `saving_adjusted: -25.66p`, GH#3894). Both are still current on `main`. Separately, the battery-value adjustment bills the SoC **level** at day-end against the counterfactual baseline rather than the **change** in SoC over the day — a day that force-exports through midnight can be energy-neutral yet still get charged the full baseline-vs-actual SoC gap as if it were lost value. | none | -| GivTCP REST (`inverter.py`) | **Structurally stale as written, and left here for the mechanism only:** PR #4864 moved REST handling out into `givtcp.py`/`givtcp_rest.py`, so `update_status()` (now `inverter.py:1333`) no longer reads `Power.Power` at `inverter.py:1435` at all. What survives is the same trap one layer up - the component's auto-config claims the power keys unless `givtcp_rest_power_ignore` is set, and it now logs an Info line when you opt out (`givtcp.py:766-767`). PR #4959 also lets an apps.yaml-named energy sensor win over auto-configuration for the history-read keys. Historically, with `givtcp_rest` configured `update_status()` read PV/Grid/Load power straight out of the REST `Power.Power` block; the apps.yaml entity lists - including any `0` placeholders put there deliberately to zero a duplicate reading - are only consulted on the non-REST `else` branch, and `execute.py` then sums every inverter's REST readings. On a hybrid + AIO pair that presents as the AIO's hybrid-fed PV port counted as solar at night, and both units' shared-CT grid readings summed (~7.1 kW shown for ~3.6 kW of real export). The per-inverter `givtcp_rest_power_ignore: true` restores the apps.yaml lists and is already documented in `docs/apps-yaml.md` - the gap is that nothing warns when a `0` placeholder is silently bypassed (GH#4883). Two v9.0.x notes since the refactor. **GH#4993 (still live on main, b8996659):** the component's required `rest_urls` arg is resolved straight from `givtcp_rest` (`components.py`) and the stock `config/apps.yaml` ships `givtcp_rest:` uncommented with example URLs, while the registry's `"inverter": True` flag is documentation-only and nothing reads it - so every install using the template, GivEnergy or not, starts a GivTCP REST component that fails discovery (`num_entities: 0`), hammers two dead URLs and shows "GivTCP REST: Error / Never updated" in the Components panel. Discovered inverters stay at zero so `automatic_config()` skips and no Predbat args get hijacked - the damage is log spam and panel error, not wrong control. Workaround: comment out the whole `givtcp_rest:` block, and any uncommented `givtcp_automatic:` line with it, because leaving that set while `givtcp_rest` is absent trips the `Warn: Skipping GivTCP REST interface, missing required configuration: givtcp_rest` gate. **GH#4984 (fixed by the refactor, keep for older logs):** `Warn: Inverter N REST failed to setDischargeRate to got ` where *got equals the requested rate* is a zero verification tolerance, not a failed write - pre-#4864 the read-back tolerance was sized from `battery_rate_max_discharge`/`charge`, which an `inverter_limit_discharge`/`inverter_limit_charge` API override of 0 (users reach for it as "manual rate 0"; it is in `CONFIG_API_OVERRIDE`) drives to 0 against a strict `<`, so an exact read-back fails, burns the 5 `INVERTER_MAX_RETRY_REST` posts (~2s apart, the 4-5 runs users see in the GivTCP log) and flags the whole run "Demand with Errors". Only fires when something else set the real rate away from the target first, otherwise `adjust_discharge_rate`'s change-gate skips the write. The refactor re-sized the tolerance from `write_tolerance_watts()` - the max rate GivTCP itself reports (`Invertor_Max_Bat_Rate`), falling back to 2600W - and the entity write path uses `<=`, so an exact match passes; the charge-rate variant (#2882, still open) is probably this same trap on an older version. **GH#5101 (still live):** `initialize()` builds one `InverterRestState(id=n, rest_api=url)` per `givtcp_rest` entry with no URL-shape validation (`givtcp.py`), so a YAML mapping indented into the list during an apps.yaml edit becomes `rest_api=` and `read_data()`'s first line (`url = inverter.rest_api + "/" + api`, `givtcp_rest.py`) raises `TypeError: unsupported operand type(s) for +: 'dict' and 'str'`, caught by `ComponentBase.start()` and logged with no index or URL. Signature: the `discovered N inverter(s) from M configured REST endpoint(s)` line never appears anywhere in the log (discovery settles only after the poll loop completes), and the malformed entry itself produces no `Errno 22` retries because it dies before `get_data()` — valid-but-dead string entries ahead of it produce the full ladder, so Errno-22 count = 5 attempts × completed poll cycles. The component never reports started; the main loop keeps working off the user's own apps.yaml entities. The whole class is new with the v9 component — v8.55 read `givtcp_rest` with `index=self.id`, so stray extra entries were inert on single-inverter installs. Also from GH#5099: nothing in the repo sets `battery_calibration` automatically — the component publishes the sensor from GivTCP's own detection but expects an apps.yaml mapping — and Predbat reads the arg's state as calibrating only when it is literally `on`/`On`/`true`/`True` (`inverter.py`), while GivTCP's detection is `Battery_Calibration != "Off"` (`givtcp_rest.py`); anyone wiring a multi-valued select into `battery_calibration` must reconcile the two. and because `refresh_config()` re-samples the arg every cycle (the objects persist since PR #5126, but the per-cycle re-read survives — see the args trap in "Traps when investigating"), `in_calibration` is re-sampled per cycle, not init-only. Another manual-vs-auto split (GH#5132): the `discharge_target_soc` write is the force-export hardware backstop - always `int(self.reserve_percent)`, never the plan's per-window export limit (that floor is software-only in `execute.py`) - and the #4517 unsupported-model guard (`DISCHARGE_TARGET_UNSUPPORTED_MODELS`/`_TYPES`, givtcp.py) runs only in the **auto-config** path, where it stops Predbat publishing the entity at all. A user on the **manual** entity-list path gets the entity from the GivTCP add-on instead, so the guard never engages: on an AIO whose add-on write handler for `number.givtcp_*_discharge_target_soc_1` is a silent no-op (diagnosed live - manual `number.set_value` did nothing and produced zero GivTCP log lines), Predbat warned `didn't complete got 4.0` every cycle for 14h. The write is skipped when the arg is absent by design, so **removing `discharge_target_soc` from apps.yaml is the first workaround on any version and either path**. Setup tell: manual-path apps.yaml lists `discharge_target_soc: - number.givtcp_*`; auto-config users don't. (Version trap from the same issue: the issue body said v8.54.3, the attached log's own version lines said v9.0.3 - the log wins.) **GH#5178 (enhancement, live on main):** whenever `givtcp_rest` is configured the component publishes its **full** entity surface unconditionally — ~44 entities per inverter — regardless of whether the GivTCP HA integration already provides equivalents; deliberate, so REST-only setups have entities at all. A "Predbat created N duplicate GivTCP entities" report maps here and is a feature request: the count is checkable as (GIVTCP_SENSORS + GIVTCP_CONTROLS − per-fleet withheld) × inverter count (44×3 = 132 in #5178, matched), the user-side workaround is an HA recorder exclusion glob `*.predbat_givtcp_*` plus hiding (safe — the four history keys keep the user's own sensors, USER_WINS), and `docs/components.md` has no GivTCP REST section yet. Related live gap (code-read, #5178's reporter's debug yaml): `automatic_config()` claims `discharge_target_soc` gated on rest_v3 alone (`givtcp.py`), while `publish_data()` additionally withholds the entity for `DISCHARGE_TARGET_UNSUPPORTED_MODELS`/`_TYPES` — on an unsupported-model fleet the arg points at `number.predbat_givtcp_N_discharge_target_soc` entities that are never created; benign so far (inverter.py leaves an unreadable discharge target alone, `No current discharge target to read`), fix shape = consult the same unsupported-model set or `published_discovery` like the discovery keys do. **GH#5209 (fixed in PR #5216; mechanism kept for pre-#5216 logs):** pre-#5216 `automatic_config()` filled every per-inverter arg list **positionally** from `discovered`, so an endpoint down at the discovery pass shifted every later endpoint down one slot, and `rediscover()`'s deliberate append-only recovery turned the shift into a rotation on recovery — recovery did not heal. Writes route by the index embedded in the entity id (`_parse_entity`), so the misroute reaches hardware; log signature `Inverter 1 current charge rate is 3000W and new target is 2600W` paired with `GivTCP: Inverter 2 set charge rate 2600 via REST successful` alternating each cycle, plus `battery_calibration ... has 2 entries, expected 3` / `Out of range index 2` for claimed-but-user-unconfigured keys (`_keep_configured_tail()` returned the short list unchanged when `_configured_value()` was None — only those keys warned while the user's own apps.yaml lists got their tail preserved). Debug-yaml diagnosis: any `predbat_givtcp__*` entry whose list position is not ``. Workarounds at any version: restart Predbat once every endpoint answers, or `givtcp_automatic: False`. Design note on the fix: at startup a dead leading endpoint cannot be told apart from a leading placeholder URL, and #5216 chooses identity-by-index — `givtcp_rest: [dead, live]` becomes two inverters, not one (`num_inverters` cannot arbitrate; the GivTCP template tells REST users to delete it). **One design rule left by #5216's own development:** during the PR's cleanup it was found that the re-probe adoption pass judged every `published_discovery` gate before the adopted endpoint had published anything — `run()` calls `publish_data()`, then `rediscover()` appends the newly answering index, then `automatic_config()` re-runs because `discovered != configured_for`, so each `all(key in self.published_discovery.get(n, set()) for n in discovered)` gate failed on the adoption cycle and discovery-gated keys (`soc_max`, `battery_calibration`, `inverter_limit`, …) were skipped or handed back while the gates reading `rest_data` directly (`rest_v3`, `pause_mode_supported`, …) judged correctly. The merged PR re-publishes after a fleet-growing `rediscover()` (`givtcp.py`, the `if rediscover:` block) and adds a claim hand-back (`self.claimed_from`), and that ordering is the rule for anyone touching the path: **any gate over `published_discovery` must not be judged in the cycle where `rediscover()` appended an index.** On pre-#5216 builds (no hand-back, no wait) a *tail* endpoint adopted late never has its discovery keys claimed at all — suspected, unverified against a build without the PR — so a "late GivTCP inverter shows 8 kWh / no calibration sensor until restart" report on an older build maps here. **A missing inverter-details block silently stales the published sensors and fakes a clock-skew alarm (GH#5334, code-read on main 57ec7bf1):** `inverter_details()` (givtcp_rest.py) has no log and returns `{}` when `Invertor_Details` is empty (normal on v3) and `raw.invertor.serial_number` is null, which is exactly the Gateway/EMS shape upstream omits from `/readData` (britkat1980/giv_tcp#597 is the upstream fix, unmerged at triage time); `publish_data()` then gates each detail-block sensor on presence — `inverter_time`, `soc_max`, `battery_rate_max`, `inverter_limit` — so HA keeps the entity's *last* state frozen and stale looks identical to live, while `serial_number` itself is published unconditionally: **the discriminating tell is the serial sensor reading unknown while inverter_time stays frozen**. Predbat reads that frozen `inverter_time` and `check_clock_skew()` (inverter.py) crosses the ≥30-min restart threshold every cycle — repeating `Warn: Inverter time is , Predbat computer time , this is minutes skewed` and `Warn: Inverter control auto restart trigger: Clock skew >= 30 minutes command []` (empty `command` = no `auto_restart` configured) with the inverter clock actually fine — not the clock, not HA time sync. Fix shapes discussed, maintainer's call: a per-cycle warn when the details block resolves empty; a fallback lookup over top-level dict blocks for `Invertor_Serial_Number`/`Invertor_Time` (unambiguous only at exactly one candidate); or the cheapest anti-misdiagnosis — publish an explicit `unknown`/`unavailable` `inverter_time` instead of skipping, which `Inverter` already treats as "no reading" and skips skew detection for without a restart trigger. Any "Predbat-published entity looks frozen but claims to be live" report on a REST-backed component is a candidate for this same `if value:` publish-gate class. | `inverter` | +| GivTCP REST (`inverter.py`) | **Structurally stale as written, and left here for the mechanism only:** PR #4864 moved REST handling out into `givtcp.py`/`givtcp_rest.py`, so `update_status()` (now `inverter.py:1333`) no longer reads `Power.Power` at `inverter.py:1435` at all. What survives is the same trap one layer up - the component's auto-config claims the power keys unless `givtcp_rest_power_ignore` is set, and it now logs an Info line when you opt out (`givtcp.py:766-767`). PR #4959 also lets an apps.yaml-named energy sensor win over auto-configuration for the history-read keys. Historically, with `givtcp_rest` configured `update_status()` read PV/Grid/Load power straight out of the REST `Power.Power` block; the apps.yaml entity lists - including any `0` placeholders put there deliberately to zero a duplicate reading - are only consulted on the non-REST `else` branch, and `execute.py` then sums every inverter's REST readings. On a hybrid + AIO pair that presents as the AIO's hybrid-fed PV port counted as solar at night, and both units' shared-CT grid readings summed (~7.1 kW shown for ~3.6 kW of real export). The per-inverter `givtcp_rest_power_ignore: true` restores the apps.yaml lists and is already documented in `docs/apps-yaml.md` - the gap is that nothing warns when a `0` placeholder is silently bypassed (GH#4883). Two v9.0.x notes since the refactor. **GH#4993 (still live on main, b8996659):** the component's required `rest_urls` arg is resolved straight from `givtcp_rest` (`components.py`) and the stock `config/apps.yaml` ships `givtcp_rest:` uncommented with example URLs, while the registry's `"inverter": True` flag is documentation-only and nothing reads it - so every install using the template, GivEnergy or not, starts a GivTCP REST component that fails discovery (`num_entities: 0`), hammers two dead URLs and shows "GivTCP REST: Error / Never updated" in the Components panel. Discovered inverters stay at zero so `automatic_config()` skips and no Predbat args get hijacked - the damage is log spam and panel error, not wrong control. Workaround: comment out the whole `givtcp_rest:` block, and any uncommented `givtcp_automatic:` line with it, because leaving that set while `givtcp_rest` is absent trips the `Warn: Skipping GivTCP REST interface, missing required configuration: givtcp_rest` gate. **GH#4984 (fixed by the refactor, keep for older logs):** `Warn: Inverter N REST failed to setDischargeRate to got ` where *got equals the requested rate* is a zero verification tolerance, not a failed write - pre-#4864 the read-back tolerance was sized from `battery_rate_max_discharge`/`charge`, which an `inverter_limit_discharge`/`inverter_limit_charge` API override of 0 (users reach for it as "manual rate 0"; it is in `CONFIG_API_OVERRIDE`) drives to 0 against a strict `<`, so an exact read-back fails, burns the 5 `INVERTER_MAX_RETRY_REST` posts (~2s apart, the 4-5 runs users see in the GivTCP log) and flags the whole run "Demand with Errors". Only fires when something else set the real rate away from the target first, otherwise `adjust_discharge_rate`'s change-gate skips the write. The refactor re-sized the tolerance from `write_tolerance_watts()` - the max rate GivTCP itself reports (`Invertor_Max_Bat_Rate`), falling back to 2600W - and the entity write path uses `<=`, so an exact match passes; the charge-rate variant #2882 (still open) does not match this trap's signature — it is a ~1 kW read-back miss (3600 written, 2560 got), not got==requested — and maps instead to the GH#5386 family appended below. **GH#5101 (still live):** `initialize()` builds one `InverterRestState(id=n, rest_api=url)` per `givtcp_rest` entry with no URL-shape validation (`givtcp.py`), so a YAML mapping indented into the list during an apps.yaml edit becomes `rest_api=` and `read_data()`'s first line (`url = inverter.rest_api + "/" + api`, `givtcp_rest.py`) raises `TypeError: unsupported operand type(s) for +: 'dict' and 'str'`, caught by `ComponentBase.start()` and logged with no index or URL. Signature: the `discovered N inverter(s) from M configured REST endpoint(s)` line never appears anywhere in the log (discovery settles only after the poll loop completes), and the malformed entry itself produces no `Errno 22` retries because it dies before `get_data()` — valid-but-dead string entries ahead of it produce the full ladder, so Errno-22 count = 5 attempts × completed poll cycles. The component never reports started; the main loop keeps working off the user's own apps.yaml entities. The whole class is new with the v9 component — v8.55 read `givtcp_rest` with `index=self.id`, so stray extra entries were inert on single-inverter installs. Also from GH#5099: nothing in the repo sets `battery_calibration` automatically — the component publishes the sensor from GivTCP's own detection but expects an apps.yaml mapping — and Predbat reads the arg's state as calibrating only when it is literally `on`/`On`/`true`/`True` (`inverter.py`), while GivTCP's detection is `Battery_Calibration != "Off"` (`givtcp_rest.py`); anyone wiring a multi-valued select into `battery_calibration` must reconcile the two. and because `refresh_config()` re-samples the arg every cycle (the objects persist since PR #5126, but the per-cycle re-read survives — see the args trap in "Traps when investigating"), `in_calibration` is re-sampled per cycle, not init-only. Another manual-vs-auto split (GH#5132): the `discharge_target_soc` write is the force-export hardware backstop - always `int(self.reserve_percent)`, never the plan's per-window export limit (that floor is software-only in `execute.py`) - and the #4517 unsupported-model guard (`DISCHARGE_TARGET_UNSUPPORTED_MODELS`/`_TYPES`, givtcp.py) runs only in the **auto-config** path, where it stops Predbat publishing the entity at all. A user on the **manual** entity-list path gets the entity from the GivTCP add-on instead, so the guard never engages: on an AIO whose add-on write handler for `number.givtcp_*_discharge_target_soc_1` is a silent no-op (diagnosed live - manual `number.set_value` did nothing and produced zero GivTCP log lines), Predbat warned `didn't complete got 4.0` every cycle for 14h. The write is skipped when the arg is absent by design, so **removing `discharge_target_soc` from apps.yaml is the first workaround on any version and either path**. Setup tell: manual-path apps.yaml lists `discharge_target_soc: - number.givtcp_*`; auto-config users don't. (Version trap from the same issue: the issue body said v8.54.3, the attached log's own version lines said v9.0.3 - the log wins.) **GH#5178 (enhancement, live on main):** whenever `givtcp_rest` is configured the component publishes its **full** entity surface unconditionally — ~44 entities per inverter — regardless of whether the GivTCP HA integration already provides equivalents; deliberate, so REST-only setups have entities at all. A "Predbat created N duplicate GivTCP entities" report maps here and is a feature request: the count is checkable as (GIVTCP_SENSORS + GIVTCP_CONTROLS − per-fleet withheld) × inverter count (44×3 = 132 in #5178, matched), the user-side workaround is an HA recorder exclusion glob `*.predbat_givtcp_*` plus hiding (safe — the four history keys keep the user's own sensors, USER_WINS), and `docs/components.md` has no GivTCP REST section yet. Related live gap (code-read, #5178's reporter's debug yaml): `automatic_config()` claims `discharge_target_soc` gated on rest_v3 alone (`givtcp.py`), while `publish_data()` additionally withholds the entity for `DISCHARGE_TARGET_UNSUPPORTED_MODELS`/`_TYPES` — on an unsupported-model fleet the arg points at `number.predbat_givtcp_N_discharge_target_soc` entities that are never created; benign so far (inverter.py leaves an unreadable discharge target alone, `No current discharge target to read`), fix shape = consult the same unsupported-model set or `published_discovery` like the discovery keys do. **GH#5209 (fixed in PR #5216; mechanism kept for pre-#5216 logs):** pre-#5216 `automatic_config()` filled every per-inverter arg list **positionally** from `discovered`, so an endpoint down at the discovery pass shifted every later endpoint down one slot, and `rediscover()`'s deliberate append-only recovery turned the shift into a rotation on recovery — recovery did not heal. Writes route by the index embedded in the entity id (`_parse_entity`), so the misroute reaches hardware; log signature `Inverter 1 current charge rate is 3000W and new target is 2600W` paired with `GivTCP: Inverter 2 set charge rate 2600 via REST successful` alternating each cycle, plus `battery_calibration ... has 2 entries, expected 3` / `Out of range index 2` for claimed-but-user-unconfigured keys (`_keep_configured_tail()` returned the short list unchanged when `_configured_value()` was None — only those keys warned while the user's own apps.yaml lists got their tail preserved). Debug-yaml diagnosis: any `predbat_givtcp__*` entry whose list position is not ``. Workarounds at any version: restart Predbat once every endpoint answers, or `givtcp_automatic: False`. Design note on the fix: at startup a dead leading endpoint cannot be told apart from a leading placeholder URL, and #5216 chooses identity-by-index — `givtcp_rest: [dead, live]` becomes two inverters, not one (`num_inverters` cannot arbitrate; the GivTCP template tells REST users to delete it). **One design rule left by #5216's own development:** during the PR's cleanup it was found that the re-probe adoption pass judged every `published_discovery` gate before the adopted endpoint had published anything — `run()` calls `publish_data()`, then `rediscover()` appends the newly answering index, then `automatic_config()` re-runs because `discovered != configured_for`, so each `all(key in self.published_discovery.get(n, set()) for n in discovered)` gate failed on the adoption cycle and discovery-gated keys (`soc_max`, `battery_calibration`, `inverter_limit`, …) were skipped or handed back while the gates reading `rest_data` directly (`rest_v3`, `pause_mode_supported`, …) judged correctly. The merged PR re-publishes after a fleet-growing `rediscover()` (`givtcp.py`, the `if rediscover:` block) and adds a claim hand-back (`self.claimed_from`), and that ordering is the rule for anyone touching the path: **any gate over `published_discovery` must not be judged in the cycle where `rediscover()` appended an index.** On pre-#5216 builds (no hand-back, no wait) a *tail* endpoint adopted late never has its discovery keys claimed at all — suspected, unverified against a build without the PR — so a "late GivTCP inverter shows 8 kWh / no calibration sensor until restart" report on an older build maps here. **A missing inverter-details block silently stales the published sensors and fakes a clock-skew alarm (GH#5334, code-read on main 57ec7bf1):** `inverter_details()` (givtcp_rest.py) has no log and returns `{}` when `Invertor_Details` is empty (normal on v3) and `raw.invertor.serial_number` is null, which is exactly the Gateway/EMS shape upstream omits from `/readData` (britkat1980/giv_tcp#597 is the upstream fix, unmerged at triage time); `publish_data()` then gates each detail-block sensor on presence — `inverter_time`, `soc_max`, `battery_rate_max`, `inverter_limit` — so HA keeps the entity's *last* state frozen and stale looks identical to live, while `serial_number` itself is published unconditionally: **the discriminating tell is the serial sensor reading unknown while inverter_time stays frozen**. Predbat reads that frozen `inverter_time` and `check_clock_skew()` (inverter.py) crosses the ≥30-min restart threshold every cycle — repeating `Warn: Inverter time is , Predbat computer time , this is minutes skewed` and `Warn: Inverter control auto restart trigger: Clock skew >= 30 minutes command []` (empty `command` = no `auto_restart` configured) with the inverter clock actually fine — not the clock, not HA time sync. Fix shapes discussed, maintainer's call: a per-cycle warn when the details block resolves empty; a fallback lookup over top-level dict blocks for `Invertor_Serial_Number`/`Invertor_Time` (unambiguous only at exactly one candidate); or the cheapest anti-misdiagnosis — publish an explicit `unknown`/`unavailable` `inverter_time` instead of skipping, which `Inverter` already treats as "no reading" and skips skew detection for without a restart trigger. Any "Predbat-published entity looks frozen but claims to be live" report on a REST-backed component is a candidate for this same `if value:` publish-gate class. **GH#5386 (2026-10-04, open) — a read-back miss larger than the REST verify tolerance never converges:** the REST verify compares within `write_tolerance_watts()/12` (charge) and `/25` (discharge), where `write_tolerance_watts()` is the max rate GivTCP itself reports (`Invertor_Max_Bat_Rate`, 2600 W fallback); the reporter's 6000 W write read back 5100 while GivTCP's own log said the POST succeeded, so every ~2s retry and every cycle re-warns, and the entity-level deadband (`rate_tolerances()`'s one-step `fuzzy_below`) cannot cover a ~9-step miss either. Predbat's REST mirror entities **are** the read-back, so the debug yaml alone (mirror state vs the target) proves non-convergence without GivTCP logs. Why the register sits at 5100 is open — a stale cached Control read vs a firmware/BMS cap (the mirror's own max attribute says 6000, inconsistent with a pure cap) — and the suspected fix shape (accept GivTCP's POST success, a Solis-style `verify_settle_seconds` re-read, or sizing the tolerance to the firmware cap) is untested. Family so far: #5324 (exactly one rounding step; fixed by PR #5348's `fuzzy_below`), #2882 (~1 kW miss, open), #5386 (900 W miss, open). | `inverter` | | Compare (`compare.py`) | `apply_hardware_overrides()` overrode four attributes and not `battery_rate_max_export`, which is the one the export prediction path actually uses (`prediction.py:906`), so an `override_battery_rate_max_discharge_kw` left force-export slots pinned at the real hardware rate - probe-verified against the reporter's debug yaml, fixed in PR #4897 with an explicit fifth key (GH#4895). `battery_rate_max_export` itself is `min(inverter_limit_export, battery_rate_max_raw)` (`inverter.py:431`), and the prediction still clamps export draw at the un-overridden `inverter_limit`, so a "model a bigger inverter" scenario needs `override_inverter_limit_kw` as well or it is inert. A negative "True cost" alongside zero import is **not** phantom export revenue: compare reports `metric_end - metric_start`, and `compute_metric()` credits the end-of-scenario battery at the *replacement import* rate (`plan.py:1769`), deliberately ignoring the export rate - so a scenario ending 25 kWh fuller on free PV scores negative by construction. Likewise a non-zero Export column under `rates_export: 0` is forced PV-clipping export, not a planned export slot; banning export outright is a missing feature (GH#2446), not a compare bug (GH#4881, confirmed twice, the second time against the reporter's own detailed plans). GH#2033 (bumped 2026-09-18, code-verified): the table's Import and Export columns are whole-horizon totals, not plan-slot sums — `run_prediction()` accumulates `import_kwh_battery`/`import_kwh_house`/`export_kwh` for every modelled minute of grid flow and seeds `export_kwh` from `export_today_now`, so PV export outside slots, scenario-plan arbitrage and today-so-far export all count, and a planned discharge slot exports regardless of load vs solar. Deliberate: compare/annual pass `include_manual_api=False` to `basic_rates()` so live manual_api overrides never leak into simulated tariffs, while a tariff's own `rates_import`/`rates_export` replaces the live rates wholesale (no `prev` passed) — per-tariff `rates_*_override` blocks are the workaround; and #2033's "identical friendly names" premise is contradicted by the code (`"Compare " + name`), so duplicate names likely come from duplicate `name:` values in `compare_list`. GH#5133 (fixed on **open PR #5143, unmerged** - keep the mechanism for logs on stock main): `publish_data()` built the entity id by plain concatenation with no slug sanitising, so a `compare_list` id containing `/` produced `POST /api/states/predbat.compare_tariff_IGO/Prime` - and that 404 is the route not matching, not a missing entity. The warns arrive **in pairs** - `set_state()` does the POST and then a read-back `update_state()` GET, and both warn - so two identical 404 warns per cycle anywhere on `/api/states/...` means one bad write, not two problems. And `comparisons.yaml` was never pruned against `compare_list`, so a renamed or deleted tariff id kept republishing for as long as the file survived: the bad id existed only in the persisted yaml (the reporter's current apps.yaml looked fine), so grepping the attached config proves nothing - ask for `comparisons.yaml`, delete the stale entry and restart. The PR adds `tariff_entity_slug()` and `prune_comparisons()`. From the same PR's review (still unmerged, so observations about stock main): the tree has four ad-hoc entity-id slugify helpers with different semantics - `gateway.py`'s `_ev_suffix` drops non-`[a-z0-9_]` characters, `sigenergy.py`'s `_system_slug` folds only `-`, web.py's `re.sub` neither collapses runs nor strips, the PR's `tariff_entity_slug` collapses and strips - and `dashboard_item()` validates nothing, so on any "invalid entity id / 404 on publish" report check which slugify helper built the id; none caps length, and HA rejects object ids over 64 characters. A history lookup that passes an unknown entity id returns `None` silently - a wasted round-trip, **not** a 404 - so don't accept "it 404s" about the history path. With `compare_list` commented out, stored tariffs are republished **once per Predbat start** (only `load_yaml` reaches `publish_data` without the `compare_list` early return), so the tell for "sensors for a tariff I deleted keep coming back" is a restart, not the 5-minute cycle; and stored compare results are looked up by the raw `compare_list` id everywhere (`run_all`'s prior-SoC carry-forward, `select_best`, web.py's Compare page), so a fix that normalises tariff identity must *move* the stored entry to the configured id - keeping the old key preserved nothing, because nothing looked it up under the new id. **GH#5180 (code-read on main, still live):** the /compare tab's `Existing` badge is broken by any event stamped into the live rate tables at the moment compare runs — `Compare.run_all()` snapshots its baseline from the live `self.pb.rate_import`/`rate_export` (`compare.py:598-599`), which by then carry the IO/saving/free/Axle stamps, while `fetch_rates()` refetches the compared tariff's plain rate card and never re-applies any event — so the minute compare differs at every event minute and `result["existing_tariff"]` goes False, withholding the badge (web.py) for that whole run. Compare fires daily at midnight or on the `compare_active` switch, so an event active at that moment is enough. Cost figures are unaffected: every scenario runs on its own refetched plain card. In-tree fix baselines already exist: `rate_import_no_io` (also excludes manual overrides — but the compared refetch never applies manual rates either, so override users already fail this comparison today; the new base makes that consistent rather than opening a gap) and `rate_export_base` (comment says saving sessions and overrides stay out of it). `test_compare` has zero setup for saving/free/IO slots and never asserts `existing_tariff` — a green compare test says nothing about the event-active path. Suspected, not verified: the `run_all()` finally block restores everything it saved **except** `rate_import`/`rate_export`, so the live tables may hold the last compared tariff's rates until the next `fetch_sensor_data()` rebuild — if "the live plan used a different tariff's rates right after running a comparison" is reported, start there. | `compare` | -| Car charging (`plan.py`, `fetch.py`, `execute.py`) | `car_charge_slot_kwh()` returns the raw slot kWh with no limit awareness and feeds the HTML Car kWh column, the JSON plan, the status sentence and the `car_charging_slot` attributes, while `prediction.py:799` and `prediction_kernel.cpp:827` clamp unconditionally - so "car kWh is displayed but the cost never moves" is a display-vs-model disagreement, not a planning bug. The *rate* side of beyond-cap IOG energy is already handled (`car_charge_slot_rate`/`car_rate_premium`); only the volume side is untrimmed, and `octopus_intelligent_consider_full` (default False, expert-hidden) is the only thing that trims it (GH#4888). The manual car SoC is prediction-fed and never a measurement: `predbat.py` exposes `car_charging_soc_next`, which `prediction.py` sets from the modelled first-step car SoC, so it never resets on its own (GH#4889). `update_car_charging_power()` (`execute.py`) sums *every* entity listed under `car_charging_power` - that list describes chargers, not cars, so two per-car template sensors reading one shared charger are added together; the summing is confirmed, the reporter's doubled figure was not (GH#4879). Per-car items are read by **suffixed name**, not index-sliced: `car_charging_rate` is read as `float(get_arg("car_charging_rate" + postfix))` (`fetch.py`), so an apps.yaml **list** for it is unsupported — and when the item is fresh (the HA input_number does not exist yet) `load_user_config()` seeds `item["default"] = self.args[name]` with no type validation (`userinterface.py`), the invalid seed survives, and `float(list)` raises `TypeError` every main-loop cycle while the car falls back to 7.4 kW. The list-vs-postfix asymmetry in the same table is the trap: `car_charging_limit`/`car_charging_battery_size` *are* read with `index=car_n` and so do accept lists. The workaround is per-car keys — `car_charging_rate: 11.0` plus `car_charging_rate_1: ...` (or the HA input_numbers) (GH#5366, probe-reproduced; suspected but not reproduced: a comma-string value `car_charging_rate: "11.0, 11.0"` hits the same `float()` on a truly fresh item). The `car_charging_slots` refresh gate that checked only `car_charging_planned` now also checks `car_charging_now` (`fetch.py:1337`, GH#4795 / PR #4796). **GH#4967 is now fixed (PR #4971, commit `4a045e07`) - the mechanism is kept for anyone reading a pre-#4971 log.** The prediction used to clamp modelled car load at the real `car_charging_limit` unconditionally, via an identical line in both `prediction.py` and `prediction_kernel.cpp`, with the discharge hold below each gated on `car_load_scale > 0` - so once the modelled car filled part-way through a slot the hold released for the rest of the window, regardless of the `octopus_intelligent_consider_full` switch, which never reached `predict()` at all (its only live uses were the read, the `False` default and the gated trim inside `load_octopus_slots()`). The fix routes the prediction through a model-facing `car_charging_limit_model` instead: with `octopus_intelligent_consider_full` off (the default) cars carrying IOG slots get `CAR_CHARGING_LIMIT_UNCAPPED`, so the fill clamp is deliberately inert and the hold releases only when the modelled car actually fills, while the real `car_charging_limit` is left untouched for `execute.py`'s "car already charged" decision, `plan_car_charging()` and `load_octopus_slots()` (it is taken **by reference** by `Prediction`, which is why the fix assigns a separate value rather than mutating it). Both engines take the model limit from the same attribute so they cannot diverge here, `update_car_manual_soc()` caps the manual car SoC write-back at the real limit because the modelled SoC can now overshoot, and `annual.py`/`compare.py` reset the attribute so an override cannot leak into their re-planning; `test_multi_car_iog.py` now asserts the switch's effect on the prediction. One interaction to re-check when the GH#4952 Ohme `watts x hours` overestimate is fixed: with consider_full off the model limit is uncapped for IOG-slot cars, so the fill clamp no longer bounds modelled car load for them either. The History view's separate reconstruction of car slots from the car energy sensor had its own faked-clock bug — GH#5004, under "The shared clock" above, fixed in PR #5025 (v9.0.2). The "Hold for car" discharge hold has **two different gate conditions** (GH#5146): live execution (`execute.py`'s carHolding block) requires an active car slot with `kwh > 0`, the car not already at limit, and skips entirely while an export window is executing (`if not isExporting:`), while the plan model (`prediction.py`, `discharge_rate_now = battery_rate_min`) holds on `(car_load_scale > 0) and (not car_charging_from_battery) and set_charge_window` — no live state, no kwh check, and `set_charge_window` is the plan's charge-window flag, not "this minute is in one". So the plan can hold in slots the live code charges through and vice versa: verify a "plan says X, status says Y" report against the right gate — and since PR #5147 the plan-table display *does* take `hold_for_car` from the prediction's own record (`predict_car_hold_best`, shown for at least half a slot on Demand rows), so a display disagreement is a version question, not a gate question. The export-window exemption is deliberate and test-codified (`discharge_car_full_bat2` in `test_execute.py` asserts export + car slot + the switch off → Exporting, no pause), so it is not a #2380 regression; if the reporter has a working `car_charging_now` sensor the overlapping export window is dropped on the next replan, so *sustained* discharge means the now-slot was never created — on pre-#5245 builds a numeric power sensor read as "not charging", but since PR #5245/#5267 `car_charging_now` can *be* a charging power sensor (a number in watts counts as charging from `CAR_CHARGING_NOW_POWER_W`, `car_charging_now_value()` in `fetch.py`), so check the version before reading a power sensor as absent evidence. Read-only users cannot see the status sensor's "Hold for car" at all — `predbat.status` is forced to "Read-Only" first (`execute.py`). **"Hold for car" is the hold working as designed, and there is no native EV-SoC or price gate anywhere (GH#4572, read on main 57ec7bf1):** the hold fires when `set_charge_window` is on and `car_charging_from_battery` is off (config.py) whenever a car is charging now or inside a planned slot; it stops the battery *feeding the car* (pause mode, else rate 0) and never blocks the car charging from the grid, so "battery idle while the EV charges on PV-then-grid" is the intended steady state and the charger's start/stop is outside the hold's scope. The full car config surface (`car_charging_energy_scale`/`_threshold`/`_rate`/`_loss`/`_hold`, `car_energy_reported_load`, `_manual_soc(_kwh)`, `_plan_smart`, `_plan_max_price`, `_from_battery`, `_plan_time`) has **no parameter gating charging on the EV's SoC and no enforceable price threshold** — `car_charging_soc`/`car_charging_limit` in apps.yaml are the car's *input* sensors, not thresholds — so "charge the EV only when X" is an enhancement ask; the maintainer-side answer is external HA automation, and the in-tree tooling direction is open PR #3791 (`predbat.solar_surplus_power` = `min(max(0, grid + car_power - max(0, battery)), max(0, pv))` + `binary_sensor.predbat_force_export_slot`, sensor-only after being pared back, still unmerged with an owner review round outstanding 2026-10-02). | `octopus_*`, `car_charging` | +| Car charging (`plan.py`, `fetch.py`, `execute.py`) | `car_charge_slot_kwh()` returns the raw slot kWh with no limit awareness and feeds the HTML Car kWh column, the JSON plan, the status sentence and the `car_charging_slot` attributes, while `prediction.py:799` and `prediction_kernel.cpp:827` clamp unconditionally - so "car kWh is displayed but the cost never moves" is a display-vs-model disagreement, not a planning bug. The *rate* side of beyond-cap IOG energy is already handled (`car_charge_slot_rate`/`car_rate_premium`); only the volume side is untrimmed, and `octopus_intelligent_consider_full` (default False, expert-hidden) is the only thing that trims it (GH#4888). The manual car SoC is prediction-fed and never a measurement: `predbat.py` exposes `car_charging_soc_next`, which `prediction.py` sets from the modelled first-step car SoC, so it never resets on its own (GH#4889). `update_car_charging_power()` (`execute.py`) sums *every* entity listed under `car_charging_power` - that list describes chargers, not cars, so two per-car template sensors reading one shared charger are added together; the summing is confirmed, the reporter's doubled figure was not (GH#4879). Per-car rate config pre-#5383: `car_charging_rate` was read by suffixed name as `float(get_arg("car_charging_rate" + postfix))` (`fetch.py`), so an apps.yaml **list** raised `float(list)` — a `TypeError` every main-loop cycle while the car fell back to 7.4 kW — and a fresh item (the HA input_number not yet created) had `load_user_config()` seed `item["default"] = self.args[name]` with no type validation, letting the invalid seed survive (`userinterface.py`; a comma-string value such as `car_charging_rate: "11.0, 11.0"` hits the same `float()` — suspected, not reproduced). **Fixed in PR #5383 (merged 2026-10-04, unreleased at the time of writing, post-v9.3.5) — keep the pre-#5383 mechanism for older logs:** `car_list_arg()` (`userinterface.py`) now splits a base-key list across the per-car slots (`car_charging_rate: [11.0, 6.5]` → car 0 11.0 / car 1 6.5), a per-car suffixed scalar (`car_charging_rate_1: 3.6`) wins over its list slice, and a list on a per-car suffix is ignored in favour of the base list's slice (warned when it stands alone) — `test_fetch_config_options.py` pins all three. The list-vs-postfix asymmetry is closed: `car_charging_limit`/`car_charging_battery_size` had always accepted lists (read with `index=car_n`) and `car_charging_rate` now accepts them. The `car_charging_slots` refresh gate that checked only `car_charging_planned` now also checks `car_charging_now` (`fetch.py:1337`, GH#4795 / PR #4796). **GH#4967 is now fixed (PR #4971, commit `4a045e07`) - the mechanism is kept for anyone reading a pre-#4971 log.** The prediction used to clamp modelled car load at the real `car_charging_limit` unconditionally, via an identical line in both `prediction.py` and `prediction_kernel.cpp`, with the discharge hold below each gated on `car_load_scale > 0` - so once the modelled car filled part-way through a slot the hold released for the rest of the window, regardless of the `octopus_intelligent_consider_full` switch, which never reached `predict()` at all (its only live uses were the read, the `False` default and the gated trim inside `load_octopus_slots()`). The fix routes the prediction through a model-facing `car_charging_limit_model` instead: with `octopus_intelligent_consider_full` off (the default) cars carrying IOG slots get `CAR_CHARGING_LIMIT_UNCAPPED`, so the fill clamp is deliberately inert and the hold releases only when the modelled car actually fills, while the real `car_charging_limit` is left untouched for `execute.py`'s "car already charged" decision, `plan_car_charging()` and `load_octopus_slots()` (it is taken **by reference** by `Prediction`, which is why the fix assigns a separate value rather than mutating it). Both engines take the model limit from the same attribute so they cannot diverge here, `update_car_manual_soc()` caps the manual car SoC write-back at the real limit because the modelled SoC can now overshoot, and `annual.py`/`compare.py` reset the attribute so an override cannot leak into their re-planning; `test_multi_car_iog.py` now asserts the switch's effect on the prediction. One interaction to re-check when the GH#4952 Ohme `watts x hours` overestimate is fixed: with consider_full off the model limit is uncapped for IOG-slot cars, so the fill clamp no longer bounds modelled car load for them either. The History view's separate reconstruction of car slots from the car energy sensor had its own faked-clock bug — GH#5004, under "The shared clock" above, fixed in PR #5025 (v9.0.2). The "Hold for car" discharge hold has **two different gate conditions** (GH#5146): live execution (`execute.py`'s carHolding block) requires an active car slot with `kwh > 0`, the car not already at limit, and skips entirely while an export window is executing (`if not isExporting:`), while the plan model (`prediction.py`, `discharge_rate_now = battery_rate_min`) holds on `(car_load_scale > 0) and (not car_charging_from_battery) and set_charge_window` — no live state, no kwh check, and `set_charge_window` is the plan's charge-window flag, not "this minute is in one". So the plan can hold in slots the live code charges through and vice versa: verify a "plan says X, status says Y" report against the right gate — and since PR #5147 the plan-table display *does* take `hold_for_car` from the prediction's own record (`predict_car_hold_best`, shown for at least half a slot on Demand rows), so a display disagreement is a version question, not a gate question. The export-window exemption is deliberate and test-codified (`discharge_car_full_bat2` in `test_execute.py` asserts export + car slot + the switch off → Exporting, no pause), so it is not a #2380 regression; if the reporter has a working `car_charging_now` sensor the overlapping export window is dropped on the next replan, so *sustained* discharge means the now-slot was never created — on pre-#5245 builds a numeric power sensor read as "not charging", but since PR #5245/#5267 `car_charging_now` can *be* a charging power sensor (a number in watts counts as charging from `CAR_CHARGING_NOW_POWER_W`, `car_charging_now_value()` in `fetch.py`), so check the version before reading a power sensor as absent evidence. Read-only users cannot see the status sensor's "Hold for car" at all — `predbat.status` is forced to "Read-Only" first (`execute.py`). **"Hold for car" is the hold working as designed, and there is no native EV-SoC or price gate anywhere (GH#4572, read on main 57ec7bf1):** the hold fires when `set_charge_window` is on and `car_charging_from_battery` is off (config.py) whenever a car is charging now or inside a planned slot; it stops the battery *feeding the car* (pause mode, else rate 0) and never blocks the car charging from the grid, so "battery idle while the EV charges on PV-then-grid" is the intended steady state and the charger's start/stop is outside the hold's scope. The full car config surface (`car_charging_energy_scale`/`_threshold`/`_rate`/`_loss`/`_hold`, `car_energy_reported_load`, `_manual_soc(_kwh)`, `_plan_smart`, `_plan_max_price`, `_from_battery`, `_plan_time`) has **no parameter gating charging on the EV's SoC and no enforceable price threshold** — `car_charging_soc`/`car_charging_limit` in apps.yaml are the car's *input* sensors, not thresholds — so "charge the EV only when X" is an enhancement ask; the maintainer-side answer is external HA automation, and the in-tree tooling direction is open PR #3791 (`predbat.solar_surplus_power` = `min(max(0, grid + car_power - max(0, battery)), max(0, pv))` + `binary_sensor.predbat_force_export_slot`, sensor-only after being pared back, still unmerged with an owner review round outstanding 2026-10-02). Two slot-selection facts from October 2026: the smart car planner ranks slots by **import price only** — `plan_car_charging()` reads `window["average"]`, while each `low_rates` window already carries an `export` average (`rate_scan_window(..., alt_rates=self.rate_export)`, `fetch.py`) that only `plan_iboost_smart()` consumes — so a cheap import band overlapping daylight is mis-ranked exactly when the export price beats it, and extra car kWh served from surplus PV costs the band's export price, not its import price (GH#5384, enhancement). And a `re:` config arg matches the **first** matching entity only — `resolve_arg_re()` breaks on the first hit (`userinterface.py`) — so a `car_charging_now: re:(sensor.wallbox_portal_status_description|sensor.myenergi_zappi_[0-9a-z]+_plug_status)` alternation consults exactly one detection sensor and the other integration's is never read (GH#5390, dump-verified: the resolved args carried a single plug_status entity); an alternation does not combine sensors, name one entity per arg. | `octopus_*`, `car_charging` | | Web, MCP and Chat (`web.py`, `web_mcp.py`, `chat.py`) | `html_plan_override` responds `{"success": true}` unconditionally and discards the override write's result, and the `run_in_executor()` it goes through submits to a thread pool without ever calling `.result()` on the future (`hass.py`), so an exception while persisting a slot override is swallowed with no log line anywhere. That code path is confirmed by reading; its link to the "manual override partially ignored" report it was found under is **not** (GH#3078). The MCP OAuth endpoint advertises `client_secret_basic`, but `oauth_token`/`_handle_authorization_code` only ever read the POST body and never the `Authorization` header, so a client authenticating via HTTP Basic looks like it is missing `client_id`; and `oauth_metadata_mcp` sets `issuer` to the bare host rather than `{base_url}/mcp`, which RFC 8414 requires when the metadata is served under a path suffix (GH#4799). Chat's provider table (`PROVIDERS`, `chat.py`) has no first-class Gemini entry; `type: openai` against Google's OpenAI-compatible endpoint works for plain chat but not for tool calling, because the tool-call accumulator keys fragments solely on `fragment.get("index", 0)`, so index-less SSE fragments collapse every parallel call into one garbled slot - the sibling `reasoning_details` accumulator in the same function already handles index-less fragments - and `thought_signature` appears nowhere in `chat.py` while the assistant message is replayed verbatim, so Gemini's required signature echo-back is dropped (GH#4904, both confirmed by reading, neither tested against the live API). Also `ChatRequestError.friendly()` hardcodes OpenRouter wording as its generic fallback (`OpenRouter returned HTTP N`), and so do the 402 and 429 branches — for any provider a non-401 HTTP failure is surfaced to the user as an OpenRouter error, confirmed live in GH#5074 where a direct-OpenAI 400 reached the reporter as "OpenRouter returned HTTP 400"; read the reporter's provider entry in apps.yaml, not the message. `_stream_chunks()` always appends `/chat/completions` to the user-configurable base URL and the catalogue fetch `/models`, so there is no config-only route to another path. The header status icons read **published entities, not live power**: `get_battery_status_icon()` (`web.py`) builds the SoC+icons string from `predbat.soc_kw` and the states of `binary_sensor._charging`/`_exporting`, which `set_charge_export_status()` (`output.py`) publishes from the end of `execute_plan()` — there is no path from the icon code back to the plan, so anything the header should distinguish has to be published by execute first (GH#5125/PR #5131). One string, three render sites — the page header, the dash Status table SoC row, and `html_api_get_status`'s `battery_html` JSON — all through `get_header_html`; page CSS lives in `get_header_html()` in `web_helper.py`, not a stylesheet, and the established pattern for anything coloured is a base rule plus a `body.dark-mode .foo` override, so an inline `style=` on generated markup cannot follow dark mode. **The chat model picker's "Default (…)" row was a silent no-op after any manual pick (GH#5230, fixed in PR #5274, merged 2026-09-27 — keep the mechanism for pre-#5274 logs):** `html_chat_model()` (`web_chat.py`) used to treat the empty id the Default row posts as "clear the conversation override" but called `set_selected_model()` only behind `if model_id:` — so the per-provider *remembered* selection (which outranks the apps.yaml default in `resolve_model()`'s chain: conversation override → remembered selection → default) kept winning in both the server and the browser's `effectiveModel()` mirror, and Default was unreachable. Editing the provider's `model:` in apps.yaml did **not** work around it (the remembered choice outranks it); clearing Predbat's chat storage did. #5274 calls `set_selected_model()` unconditionally (`chat_store.py` already pops the entry on a falsy id, so the default is not pinned) and clears the JS mirror's remembered copy before the picker is redrawn, and the new tests finally cover the empty-id Default path. Two /apps editor traps (PR #5243 review, code-read on main — the PR itself merged as 6d394b7a): `resolve_arg()` resolves templates with `value.format(**self.args)` and catches **only `KeyError`**, so an unbalanced brace (`abc{def`) raises `ValueError` and a positional `{0}` raises `IndexError` — `WebInterface.resolve_value_raw()` calls it for every string containing `{` on the `/apps` page render path, so a non-credential value like that is *suspected, not reproduced* to 500 the whole page (credentials escape on the page only because `mask_secret_args()` has already turned them into `xxx`); and the editor addresses nested values by dotted path (`parse_yaml_path()` splits on `.` and `[n]`) while `render_type()` builds `data-nested-path` from the literal key, so a dict key that literally contains `.` or `[` cannot be edited correctly — the save targets the wrong node or raises. The quote-bearing-value truncation (`value="${currentValue}"` in the JS) that #5243 fixed — input values now set via `input.value` and `escapeHtml` — matters only on pre-#5243 trees. **Plugins: endpoints work, nav is hardcoded (GH#5367, code-verified on main).** Plugin `on_web_start` hooks fire in `WebInterface.start()` (`web.py`) *before* the `registered_endpoints` loop mounts routes — the plugin system is initialised early (`predbat.py`) exactly for that — so a plugin page that 404s is an `on_web_start` problem, not a routing one. But `get_header_html()` (`web_helper.py`) hardcodes the nav menu with no plugin extension point, and `registered_endpoints` is consumed only by `start()` — no page lists them — so a plugin page is reachable only by typing its URL, despite `docs/plugins.md` promising "additional web interfaces and dashboards". "My plugin's page loads but I can't get to it from the Predbat UI" → the gap is nav/UI exposure, not the endpoint mechanism; fix shape is an opt-in plugin nav API (label+path, not auto-listing `registered_endpoints` — plugins register API-only paths too), matching the existing conditional-nav precedents (the Chat link's gate, "Metrics"). | `web_*`, `web_mcp`, `web_chat` | | Self-update (`github.py`, `download.py`) | The whole update path calls the GitHub REST API **unauthenticated** - the release check (`github.py:126`) and the file listing (`download.py:79`) - so it shares the 60-requests-per-hour-per-IP unauthenticated quota with everything else behind that egress IP. Behind CGNAT, a VPN or a shared proxy that 403s even though Predbat's own volume is one check per two hours; failures are deliberately never cached, so it then retries every 5-minute cycle. The listing request's error reaches only the addon stdout via `print()`, not the HA log a user would paste, which is why the report reads as "it just won't update" (GH#4886). A phone-hotspot A/B is the cheap differential test. | none | | Cloud / divergence modelling (`fetch.py`, `plan.py`) | Until v8.55.0 `step_data_history()`'s modulation was inert for every caller but one, through operator precedence: `int(...) + 1 if flip else 0` parses as `(int(...) + 1) if flip else 0`, a constant `0` whenever `flip` is False - and `plan.py`'s PV10 call was the sole `flip=True` caller. So `metric_cloud_enable` only ever perturbed the PV10 scenario, and `metric_load_divergence` - genuinely computed from load std-dev/mean in `output.py` - did nothing at all, on defaults-on settings, for as long as the block existed (GH#4870, evaluated directly rather than inferred; fixed in `f6c925d8` and then reworked again by the envelope model in `79f28f8f`). Worth knowing when reading a pre-v8.55.0 log or replaying an older debug dump: neither knob can be the explanation there. The review of the fix also found that switching the block *on* is the risky half - an empty `pv_forecast_minute10` pins `metric_cloud_coverage` at exactly 0.5 regardless of the weather, which is every non-Solcast user. | @@ -144,7 +144,7 @@ Grep for the named symbol rather than trusting a line number. | Carbon intensity (`carbon.py`, `fetch.py`) | The Carbon Intensity API (`https://api.carbonintensity.org.uk`) publishes only ~48h ahead and often far less - on 2026-09-06 it had nothing at all past end-of-day, national and regional alike, so an `fw48h` request from *now* returned 15h. Two consequences worth separating. (a) The API spec requires `{from}` as an ISO8601 `YYYY-MM-DDThh:mmZ` datetime; `fetch_carbon_data()` passed a bare `YYYY-MM-DD` and made *two* calls, the second from now+48h, which lands past the horizon and returns a bare `{}` with HTTP 200 - logged at `Error:` severity every fetch. That is upstream truncation surfacing as a Predbat error, not a broken integration - **fixed in PR #4958**: `fetch_carbon_data()` now makes one spec-format call built from `datetime.now(timezone.utc)` (`carbon.py:49-53`), which covers the whole published horizon, and the comment there records why UTC rather than a local date. (b) `carbon_intensity.get(minute, 0)` in `prediction.py:1279/1298` scores an uncovered minute as **zero** gCO2/kWh, so the tail of the plan looks carbon-free and `best_carbon` swings wildly between cycles (410kg vs 1365kg on consecutive runs). The zero-scoring itself is still in the code, re-verified on main; what mitigates it is `carbon_replicate()` (`fetch.py:1767`, alongside `rate_replicate`), which extends the forecast forward on the 24h cycle so fewer minutes fall through to the default. Check `metric_carbon` before calling this a planning bug - at 0p/kg it only distorts reporting. | `carbon` | | Storage cache expiry (`storage.py`) | Expiry is judged by the **real clock, never Predbat's**: `load()` and `cleanup_expired()` compare the `.meta` sidecar's expiry against `datetime.now(timezone.utc)`, and the hourly cleanup **deletes** an expired file plus sidecar — so an entry whose expiry was built from `self.now_utc` is born expired once Predbat's clock and the real one drift apart (debug replay, paused HA, tests). Every `expiry=` call site on main derives from the real clock; `save_plan()` was the last exception and PR #5100 fixed it (GH#5079). This is the counterweight to the "build test dates from Predbat's clock" trap below — the discriminator is *which clock the code under test compares against*: slot/rate logic measures against `midnight_utc`/`now_utc`, storage expiry compares against the real one. Test tip from the same investigation: before chasing a (culprit, victim) pair from the fuzzer, run the victim alone **and at more than one time of day** — a test that builds its "expired" timestamp from the fixture clock brackets the day rather than failing outright (`plan_persistence` Test 5 was broken on its own). | `plan_persistence` | | Manual override selects (`userinterface.py`) | The option list for a manual override select has **three independent builders**, each ending in the same `if values not in time_values ... item["options"] = ...` tail: `manual_times()` (the time-slot selects, `manual: True`), `manual_rates()` (the rate/value selects, `manual_rate: True`) and `api_select_update()` (`manual_api` only, `api: True`; `api_select()` is its read side). A change to how the lists look or are ordered must be made in all three or the dropdowns diverge — `manual_api` is easy to miss, it lives ~100 lines above the other two and is worded differently. `manual_select()` cannot drive `manual_api`: it branches on `manual_rate` and otherwise falls through to `manual_times()`, which would render an API-command select as time slots (GH#5105, verified by reading and by driving both paths from the PR #5108 test). Two incidental details: with nothing selected `values` is `""` and that tail appends an **empty-string option** (rendered as `None` on the web config page), and nothing indexes `options` positionally anywhere — selection matches the literal `"off"` — so a reorder is safe despite `impact()` rating all three methods HIGH (they sit on `initialize`/`update_time_loop`/`run_time_loop`; the risk is cosmetic). | `manual_api` | -| Marginal rate sensors (`marginal.py`) | The `binary_sensor._marginal_rate_now__is_cheap`/`_is_moderate` flags are **plan-relative predictions, never live inverter state** (GH#5107): `calculate_marginal_costs()` injects +1/2/4/8 kWh over the next hour (`MARGINAL_TIME_OFFSETS`) into a fixed best-plan prediction and takes the raw `run_prediction()[0]` delta. Bands: cheap ≤ 1.2 × `rate_min`, moderate ≤ max(0.5 × `rate_max`, 1.5 × `rate_min`) — on a 6.9p/28.9p tariff that is a narrow 8.3p–14.4p band. The delta **excludes `compute_metric`'s battery_value and metric_keep terms**, so battery-served extra load is priced only via displaced future import/export — while the plan is overnight charge-and-export arbitrage an extra kWh costs the *export* rate in the model, so cells flip mid-band while the battery is visibly exporting, and the overnight export slots sit near the profitability threshold and enter/leave the plan between updates. The sensors only refresh at the `calculate_marginal_costs()` call site in `plan.py`'s publish path, so they can be up to `calculate_plan_every` (default 10 min) stale. Start here on any "marginal sensor looks wrong" report; Sigenergy EMS export outside the plan is invisible to them. | none | +| Marginal rate sensors (`marginal.py`) | The `binary_sensor._marginal_rate_now__is_cheap`/`_is_moderate` flags are **plan-relative predictions, never live inverter state** (GH#5107): `calculate_marginal_costs()` injects +1/2/4/8 kWh over the next hour (`MARGINAL_TIME_OFFSETS`) into a fixed best-plan prediction and takes the raw `run_prediction()[0]` delta. Bands: cheap ≤ 1.2 × `rate_min`, moderate ≤ max(0.5 × `rate_max`, 1.5 × `rate_min`) — on a 6.9p/28.9p tariff that is a narrow 8.3p–14.4p band. The delta **excludes `compute_metric`'s battery_value and metric_keep terms**, so battery-served extra load is priced only via displaced future import/export — while the plan is overnight charge-and-export arbitrage an extra kWh costs the *export* rate in the model, so cells flip mid-band while the battery is visibly exporting, and the overnight export slots sit near the profitability threshold and enter/leave the plan between updates. The sensors only refresh at the `calculate_marginal_costs()` call site in `plan.py`'s publish path, so they can be up to `calculate_plan_every` (default 10 min) stale. Start here on any "marginal sensor looks wrong" report; Sigenergy EMS export outside the plan is invisible to them. **The band stats read event-contaminated values (GH#5382, dump-verified 2026-10-04):** the free/saving/Axle loaders run before the final rate scan, so with an Octopus free session in the horizon `rate_min` is 0, `cheap_threshold = rate_min × 1.2 = 0` makes `is_cheap = (cost <= 0)` permanently false (a marginal cell is never exactly 0 — extra load is modelled as displacing stored battery) and `moderate_threshold` swallows everything; while the sensor's own published `rate_min`/`rate_max` *attributes* come from `rate_min_base`/`rate_max_base` (`marginal.py`) — band and attributes disagree, and the attributes are the base-tariff shape the band should follow (fix direction: derive the band from the base values it already publishes). | none | | Manual rates (`fetch.py` `basic_rates()`/`rate_replicate()`) | `day_of_week` rules are stamped across the **whole rate horizon**, each day against its own weekday (**fixed in PR #5173**, merged 2026-09-24, GH#5168 — previously only the explicitly modelled days were filtered and the day+2 extension write seeded whatever weekday day+2 actually was not, so the *last rule processed for a time-of-day slice won* beyond the modelled days and a weekend catch-all flattened the weekday peaks exactly 48h out, every later day then taking the previous day's pattern through `rate_replicate()`, which has no `day_of_week` notion; start there on any "next-but-one day's manual rates are wrong" report). Other pre-#5173 shapes: a midnight-spanning weekday rule never reached day+2's early hours (the wrapped half seeded day+3 instead), and yesterday was checked against *today's* weekday (`int(minute_index / 1440)` truncates negative fractions to 0). Keep the mechanism for pre-#5173 logs. A consequence worth knowing before calling anything a bug on a current build: a fully-tiled manual tariff is never `rate_replicate()`d any more, so `metric_future_rate_offset_*`, `futurerate_adjust_*` and the replicated plan markers no longer touch manual-tariff minutes — "the offset stopped applying" on a manual tariff is this. **Still live — two traps that survive the fix.** A rule ending `:59:59` leaves the **final minute of every window unwritten**: `time_string_to_stamp()` floors `06:59:59` to minute 419 and the window is `range(start, 419)`, so today that minute keeps `basic_rates()`' pre-init 0 (`for minute in range(24*60): rates[minute] = 0`) and `rate_replicate()` copies it forward (probe-verified: 0p at 06:59/08:59/16:59/20:59/23:59 each day, `rate_minmax()` min 0); symptom `Import rates: min 0` / `Export rates: min 0` from a manual tariff with no genuine zero rate; workaround: end each rule at the next rule's start (`07:00:00`); in tests, assert window interiors, not exact whole-range equality. What the planner does with a lone 0p import minute is unverified (raised on PR #5173's review thread, not fixed there). And a minute left unmatched by the day filter falls to `rate_replicate()`, which applies `metric_future_rate_offset_import`/`_export`, so a test that only matches the tariff's own rates must pin that offset to 0. The fix's tests (13/14 in `test_basic_rates.py`, plus third-day assertions in 6/7) pin `day_of_week` on every horizon day in either rule order — before #5173 they asserted only today and tomorrow, so keep the lesson: a `day_of_week` regression can land green against tests that never look past tomorrow. | `basic_rates` | | Manual load adjust (`fetch.py` `step_data_history()`) | `select.predbat_manual_load_adjust` is **additive**, never a replacement (GH#5176): `step_data_history()` adds `load_adjust.get(minute_absolute, 0) * step / plan_interval_minutes` on top of the ML/historical load (floored so total load cannot go below zero), and `output.py` mirrors it for the published `load_energy_predicted`. The docs disagreement that misled the reporter ("overwrite predicted load" in docs/manual-api.md vs "adjustments" in customisation.md) was fixed by the maintainer (`c06bff91`, 2026-09-23 — docs now say "add the adjustment amount (which can be positive or negative)"), though the plan-UI cell still says only "Adjust load (kWh)" with no sign hint. Triage method: the reporter's plan HTML marks overridden slots with the `±` symbol — that marker plus the neighbouring un-marked slots attributes "applied but additive" vs "genuinely ignored" in one glance; a +0.2 on top of a multi-kWh ML slot is invisible in the SoC drop, which is the signature of "user expected replacement semantics". `load_today_comparison` (registered) asserts the additive behaviour (+1.0 → predicted load +1.0). Recurring upstream cause: free-event load baked into history — ML training has no free-event exclusion and `load_scaling_free` only scales *forecast* event slots (GH#4942 open). The select's horizon is `MANUAL_RATE_MAX_MINUTES` (7 days, `const.py`), so same-day overrides cannot be dropped by the horizon gate. Suspected, not verified: a small (~0.1-0.2 kWh/slot) systematic offset between ML sensor values and the plan's displayed load on non-overridden slots — do not do per-slot arithmetic beyond the clean case without re-deriving it. | `load_today_comparison` | | Status health (`output.py`, `predbat.py`) | False `Status: (unhealthy)` on the web dashboard (GH#5181, **still live on main**, re-verified 2026-09-26): `predbat.status`'s `last_updated` attribute is written with `now_utc_real.strftime(TIME_FORMAT)` (`output.py:2699`; `TIME_FORMAT` is `%z`, which never emits a colon on any Python) but `is_running()` parses it with `datetime.fromisoformat()` (`predbat.py:1960`), which is colon-strict before Python 3.11 — on ≤3.10 every parse raises `ValueError` → `return False` → `web.py` renders "(unhealthy)" and the API ping returns 500. On 3.11+ the relaxed parser accepts `+0100`, which is why CI on 3.14 stays green. `load_plan()` is **not** affected: its writer uses `.isoformat()` (always colon form), so that round-trip works pre-3.11 too. Trap for the obvious `str2time()` swap: it would lose the naive-string tolerance `is_running()` deliberately keeps (an install upgraded from before the timezone-aware write persists a naive `last_updated`; `web.py:988` already parses this attribute with `str2time()` — in-tree precedent but naive-intolerant) — fix the writer to `.isoformat()` or keep the naive branch. Suspected, not verified: which shipped environments actually run ≤3.10 (no addon/Dockerfile in-repo to read); no test can settle it (Python-version dependent). | none | @@ -181,7 +181,7 @@ Grep for the named symbol rather than trusting a line number. | I set a config item above its maximum and it stays capped — no YAML workaround | An `APPS_SCHEMA` item's `max` is enforced as a hard clamp on **every** value source, not just the HA UI: the input_number branch of `userinterface.py` re-clamps after `float()` (deliberate — an apps.yaml override bypasses the HA entity's own min/max; `Warn: Config item ... clamping to ...` is the diagnostic). GH#5070 instance: `debug_history_count` max 50 while the capture path itself is uncapped (retention is `interval × count` in `_capture_debug_history()`), so widening the window was a schema change — PR #5071 (merged 2026-09-19) raised the max from 50 to 500, with the storage warning in docs/customisation.md (the top of the range is ~2.5 GB on disk at 2-5 MB per snapshot). | | Hundreds of identical `Trying to write N to X didn't complete got M` warnings overnight | Before concluding control is broken, check whether the planner even had a window in that period: the writes may be the routine idle reset — `execute.py` re-writes charge/discharge rate to max every cycle while not charging so PV can still charge (GH#5073, where the planner's own overnight windows were empty and no charging was lost — pure log noise). `write_and_poll_value()` has no failure dedup: after `INVERTER_MAX_RETRY` attempts it warns once, flags `had_errors` and clears the ledger entry, so the next cycle re-attempts from scratch — a firmware state that permanently refuses a register (SolarEdge preserve-charge holding its charge-limit register at 0 is the first confirmed instance) produces ~10 refused writes + a warning per cycle, indefinitely; any brand with a refuse-silently register maps here. Exception since PR #5297: on the `GWMQTT` (Predbat hub) type the retry policy is 3 attempts, and a same-target control failing twice degrades to one write per 5 minutes until it verifies - the "indefinitely" reading is a non-gateway-type reading only. **A stable quantised read-back below the verify tolerance never converges (GH#5324, verified on main):** the entity-write tolerance is `battery_rate_max_charge * MINUTE_WATT / 20` — 5% of the **allocated** rate, and `battery_rate_max_charge = min(inverter_limit_charge, battery_rate_max_raw)/MINUTE_WATT` (`inverter.py`) — so an `inverter_limit_charge` cap (including the API override, which is in `CONFIG_API_OVERRIDE`) below ~5 register steps makes a quantised read-back (GE: a 1300 W write sits stably at 1206) fail verify forever: 10 retries plus a full repeat every cycle, while the identical write at the uncapped rate verifies (any allocation below ~1.9 kW on a register-quantised rate is the risky band). The REST layer's tolerance for the same write is sized differently — `write_tolerance_watts()/12` from `Invertor_Max_Bat_Rate` with a 2600 W fallback (`givtcp_rest.py`) — and is immune to the cap, so two layers with different tolerances answer one write and the narrower one decides. The signature is a read-back **stable** at the same value off-target every cycle (never self-heals) - distinct from the settling bounce above, where the value sits between old and target only transiently. Workaround keeps the cap exactly: pick a register-exact override value. **Fixed in PR #5348 (v9.3.4):** `rate_tolerances()` (`inverter.py`) now widens the read-back tolerance **below** the written rate to one hardware register step — `fuzzy_below = max(fuzzy, step + 1)` where step is `INVERTER_DEF`'s `rate_step_percent_of_capacity` (GE ships 1) × `nominal_capacity` — mirroring the REST layer's step, and the GivTCP REST and Gateway MQTT paths got the same rule (`49b90af6`, `e7ca9b6a`), so a rounded-down quantised read-back within one step of the write verifies instead of restarting the retry loop; a read-back over the rate is never that rounding and keeps the plain fuzzy (one step **up** is still written down to a 0W hold), and where no `rate_step_percent_of_capacity` is declared the pre-#5348 5%-of-allocated-rate fuzziness still decides. Two case notes from that investigation: the preserve-mode hypothesis could not be verified from the dump because Predbat reads no SE battery-state entity, so the preserve flag is invisible in debug dumps; and a second execute pass ~46s after the first roughly doubled the warning volume in that log — observed only, mechanism never pinned down. `'Connection to inverter ID N failed'` on SE comes from the HA SolarEdge Modbus Multi integration, not Predbat. | | Low-rate sensors stuck on / the day rate classed as low rate while a saving session or Axle event is in the forecast | The auto threshold in `set_rate_thresholds()` (fetch.py) is computed from boost-contaminated stats. Both event loaders write the reward into the *import* rate table too, tagging the minutes `saving` in `rate_replicate`, and `rate_scan()` runs afterwards — so with `rate_low_threshold: 0` the auto branch's `rate_max - 0.5` rides on a `rate_max` that includes the boosted minutes: a +100p event on a 25.95/3.49p tariff sets the import threshold to 125.45p and every window classifies as low (GH#5050, reproduced in a test through `rate_scan` → `set_rate_thresholds` → `rate_scan_window`). `rate_average` is inflated too (30.54 vs ~17.7 true in the repro), so the *manual* multiplier path is not a clean workaround either — it must be picked against the inflated `rate_average` the log's `Import rates: min/max/average` line prints, and in the repro even `0.9` still lands above the day rate (27.49 > 25.95). The in-cycle auto-correction (`rate_import_cost_threshold = highest` after `rate_scan_window`) fixes the threshold but comes after `low_rates` was already built with the wrong one, and every same-cycle consumer reads the list — the low-rate sensors (output.py), the car plan, and `charge_window_best = clone_windows(low_rates)` when `set_charge_window` is on; the price-level search re-prices from per-minute rates so plan outcomes are usually near-correct, it is the candidate granularity that degrades. The export side is mirrored: the auto export branch cross-compares `rate_export_min` against import `rate_min`, both boost-contaminated. Fix key: both loaders already tag the boosted minutes `saving` in `rate_replicate`, so derive the threshold stats from untagged minutes — no new tracking needed; the in-tree precedent for keeping event prices out of a stat is `rate_export_max_forward`, deliberately computed from `rate_export_base` before the boosts. **The import side has the same distortion through the scan overwrite, and it *widens* the cheap band (GH#5237, code-read + `car_charging_smart` suite 11/11 on main):** the stamp precedes the scan (`load_saving_slot()` runs before `rate_scan()`), so `rate_max` includes stamped minutes and the provisional `rate_max - 0.5` threshold rises with it — then the scan overwrite (`rate_import_cost_threshold = highest`, fetch.py) admits minutes that fail the provisional but pass the final: on the #2716 arithmetic a session pushed `rate_max` 25.2→34.6p and the plain 25.2p peak entered `low_rates`. `low_rates` feeds **both** car planning and the `charge_window_best` seed (`plan.py`), so battery charge-window candidates are affected too. Fix direction already in-tree: `rate_max_base` (pre-stamp max, `fetch.py`) exists, and the export side already uses the base-tariff concept for exactly this class (`rate_export_max_forward_calc`). PR #5163 (open) rewires the stats to event-free (`rate_minmax_excluding_saving`) but does not touch the overwrite itself. Also re-verified from the same issue: the high-price charge gates (`plan.py`) skip only while `charge_limit_best[window_n] == 0` — already-limited windows are exempt, so "the gate let a high-price window through" reports must first check whether an earlier pass gave it a limit. | -| Energy totals jump or reset | Provider-side cumulative counters and any monotonic clamping. | +| Energy totals jump or reset | Provider-side cumulative counters and any monotonic clamping. **Check the counter's read pattern for a dip-and-recovery first (GH#5388, 2026-10-04, code-verified):** `minute_data()`'s smoothing branch pads a dip forward at the low value or interpolates clamped to `last_state`, but only zeroes an **upward** recovery when its per-minute step exceeds `MAX_INCREMENT` = 1.2 kWh/min (`const.py`) — so a dip held long enough that its recovery ramp is gentler than 1.2 kWh/min (gap of at least R/1.2 minutes for a recovery of R kWh) is counted in full as that day's energy (GH#5388's 181.3 kWh was exactly a recovery size), while a single-read upward jump is zeroed instead; the asymmetry is the signature. Applies to any integration binding `pv_today`/`load_today`/`import_today`/`export_today` to a cumulative counter; on SolaX the same shape was behind GH#5356/#5388 (see the SolaX row) and PR #5389 added a SolaX-side `hold_counter_dip()` — the `minute_data` asymmetry itself is unfixed and explains the symptom on other integrations. | | Works in HA, broken in Docker/standalone | `ha.py` websocket and `userinterface.py` callback return values. | | Every charge/export window executes ~N minutes late | The inverter's own clock. Compensation is manual-only via the four `inverter_clock_skew_*` settings. The warn band has moved twice: originally only ≥30 minutes was flagged (restart threshold, `INVERTER_CLOCK_SKEW_RESTART_MINUTES`), so a consistent 25-minute drift sailed under it while still logging a `difference -25.47 minutes` line every cycle (GH#4927); since PR #4991 (`87d4703a`, v9.0.2) `check_clock_skew()` also warns on a moderate 5-29 minute skew, repeated at most hourly per inverter (`INVERTER_CLOCK_SKEW_WARN_MINUTES` = 5, `INVERTER_CLOCK_SKEW_WARN_REPEAT_MINUTES` = 60), so a silent moderate skew is no longer the explanation on a current version — but "current" means v9.0.2+: a reporter on v9.0.1 or earlier is back to "only ≥30 minutes was flagged". Measured from one reporter's log: an export window written at 23:05 drew house-load import until 23:35, and a charge window enabled at 00:30 did not start until ~01:00 (GH#4927). | | GE (GivEnergy-mode) discharge window times only appear in the inverter portal once the window starts | As-coded, not a bug (GH#2649, verified on main): `adjust_force_export()` nulls the start/end times when `inv_has_ge_inverter_mode and not force_export` ("no point in changing times before we enable export", `inverter.py`); both the pre-window call (execute.py, inside the `set_window_minutes` look-ahead — it passes `False` + the plan times for GE-mode inverters) and the every-cycle outside-window clear (`adjust_force_export(False)`) therefore carry no times, and times are written only by the in-window TARGET engagement branch (`adjust_force_export(True, ...)`) — with the ~10 s-per-polled-write chain (idle start → discharge start → discharge end → enable switch → mode) plus the 30 s "Sleeping (workaround) as start/end of discharge window was just adjusted" settle per real change, so full rate can lag the window start by ~10 min including the inverter's own ramp. The pre-window TARGET programming is gated on `set_window_minutes` (= `plan_interval_minutes` since v8.27.5 — `db3b78c3`/#2865 + `91f60d51`/#2875; a hard-coded 30 on v8.24.x), and `inverter_set_charge_before: false` zeroes **both** the charge and the discharge look-ahead (`fetch.py`, `fetch_config_options()`) though docs/customisation.md describes the switch as charge-only — no discharge-only knob exists (the reporter's ask stands). A freeze-sentinel (99) export limit never engages a timed export on any version (`export_mode_of()`, `utils.py`) — no start/end writes in any cycle until a mid-session replan downgrades the 99 to a real TARGET limit; the DFS window that started 15 min late did exactly that. Discriminator for "Predbat writes X randomly" with an HA Activity screenshot: Predbat writes appear as rows marked **"triggered by action Select: Select"** (attributed to the account owning the HA token); plain "changed to X" rows matching no Predbat log line are the integration echoing inverter state — the echoed `20:58:00` was the inverter's own programmed slot, which Predbat read back as `Base export window [...]`. Per-inverter asymmetry (one unit churns, the other never changes) can also be legitimate write dedup — `adjust_force_export` writes only when that inverter's read-back differs from the plan — but an unmarked echo is the discrimination above, not dedup. | `inverter` | @@ -190,6 +190,7 @@ Grep for the named symbol rather than trusting a line number. | "Predbat Core Update shows in a group called unknown" (an HA update surface group heading) | `update.predbat_version` is a **state-machine-only entity** — published via AppDaemon `set_state` in `expose_config()` (`userinterface.py`, the `type == "update"` branch) with no entity-registry entry and no device, which is true of every Predbat entity and cannot be changed from attributes. Predbat already publishes every naming attribute it can (`title`, `friendly_name`, `entity_picture`, `installed_version`, `latest_version`, `release_url` — `config.py`), and the screen in the screenshot renders them (the bold line is `friendly_name`, the grey "Predbat v9.3.4" is `title` + `latest_version`). The literal "unknown" group heading comes from the HA surface that names a grouped update entity by its **device** and had none — Settings → Updates (`ha-config-updates.ts`) resolves device-name-or-`friendly_name` and has no literal "unknown"; the exact screen/HA version that prints "unknown" was **never pinned** (asked the reporter on GH#5374, awaiting reply) — check more-info-update.ts / ha-config-updates.ts at the reporter's HA version before treating the fallback as known. Not fixable from Predbat: no device to bind (GH#5374, enhancement). | | Charge or discharge rate pinned far below the hardware's | The rate binding in the component's auto-config, not the planner: GEC takes it from the `charge_rate` entity's `max` (GH#4908), Solis bound both directions to `max_charge_power` until PR #5308 made it publish the larger of the two per-direction limits (GH#4940), Fox sums a duplicated `batteryList` (GH#4919). Check `battery_rate_max_raw` / `soc_max` in the debug yaml against the nameplate before reading any plan code. | | Low power charging mode does nothing / still charges at full rate | Usually working as designed, so check the window times before reading any code. `find_charge_rate()` (`utils.py:1726-1728`) returns max rate whenever the PV forecast summed over the rest of the charge window exceeds `LOW_POWER_PV_THRESHOLD` - **0.1 kWh flat** (`const.py:83`), not scaled to system size. Log signature: `Low power mode: PV forecast in window X.XXkWh > 0.1kWh, default to max rate`. The guard is deliberate (PR #4373): the planner costs charge windows at full rate and low power is applied only to the final plan, so a throttled rate in a PV-overlapping window would defer the target into later grid import. The GH#4577 dawn split (`LOW_POWER_PV_LIGHT_FRACTION`, `const.py:93`, applied in `fetch.py:1863`) only protects the *pre-dawn* part of a window, so a window lying entirely in daylight is all PV-light and low power can never engage there - midday windows are expected to fail. `execute.py` recomputes the rate every 5-minute cycle while the window is active, which is why a manual charge-rate workaround gets overwritten and "it used to work" reports follow. Classify as enhancement, not defect; the two levers if a fix is wanted are scaling the threshold as the dawn split does, or costing a low-power variant in the planner. Related: GH#3311, GH#4557, GH#4699, GH#4975. **On a dawn-split window the dark half is usually frozen, not throttled (GH#5241, code-read verified against main 2026-09-26):** the optimiser's limit for the dark half comes out at the reserve — the freeze sentinel `is_freeze_charge()` compares the limit percent against — so the plan HTML renders `FrzChrg` and no charge rate is computed at all ("rate zero" is really "frozen"); low power cannot engage there both because candidate costing runs with it off (`prediction.py`'s `set_charge_low_power` gate passes only `save in ["best","best10","test"]`, and every optimiser sim — batch path included — runs with `save=None`) and because the light half then abandons low power by design through the `solar_full_rate` guard above. Split boundaries land at the PV-bucket granularity (30 min), not at the logged dawn minute; a pre-dawn half reading the reserve (e.g. `03:00–08:30 @ 14.63p 4%`) in the `Raw charge windows` log lines is the tell, and a pre-dawn-only window that the plan genuinely needs energy from still charges and throttles normally — the freeze is an economics outcome, not a split defect. Toggling `set_charge_low_power_solar_full_rate` changes only the light half's rate. **Write-path variant (GH#5182, dup of #3311 - PR #4645 still unmerged), the answer when low power computes the right rate but the inverter still charges at full power on a service/script-driven inverter:** with `output_charge_control: "power"` the dummy `charge_rate`/`discharge_rate` entities are auto-created only when `inv_output_charge_control != "power"` (`inverter.py:630`), so there is nowhere to store the computed rate; `get_current_charge_rate()` then falls back to `battery_rate_max_raw` (`inverter.py:1952`), and `adjust_charge_immediate()` fills the service template's `{power}` placeholder from it — `{power}` is always the inverter max on a no-entity inverter, and where a #3311 helper entity exists it trails every rate change by one 5-minute cycle. Mixed-fleet extension (GH#5359, probe-verified): when the fleet's args carry the percent key at all, these same read gates take the percent branch regardless of the slot's own state — see the GE Cloud row. Log signature: `Best rate: 600W` → service payload `{'value': '2500'}` → `count register writes 0`, with no `PV forecast in window` line. **The execute suite cannot see any of this (GH#5252, verified 2026-09-26):** `test_execute.py`'s mock inverter **overrides** `adjust_charge_immediate()`/`adjust_export_immediate()` (recording only the SoC/freeze arguments), so the real payload-building code — where `{power}` is filled from `get_current_charge_rate()`/`get_current_discharge_rate()` (`inverter.py`) — never executes under `execute_plan` tests, and the isolated-method tests pre-set the state so they cannot see execute_plan's ordering either; a rate-write ordering change (`72b5817a` moved the rate writes out of the execute_plan loop while the immediate calls stayed in-loop reading the stored arg) passed the whole execute suite with zero assertion edits. A green execute run is **not** evidence about the service payload — check the mock's override list before trusting it, and a fix for #5252 needs a test that runs the real `adjust_charge_immediate` against a two-cycle `execute_plan` fixture. | +| The plan tab's tooltip charge kW is far higher than the plotted `battery_power_best`, while the plan itself is right | Display-only (GH#5394, v9.3.5): the tooltip's `get_charge_rate_kw()` (`output.py`) re-runs `find_charge_rate()` with the prediction's inputs and returns there, never applying the "Clamp at inverter limit" block (`prediction.py`) that converges the DC-side charge to `min(rate, inverter_limit × inverter_loss)` through `get_total_inverted()` — the reporter's 14.4 kW request landed at 10.56 kW DC. A hybrid with PV can legitimately exceed the AC limit via `battery_rate_max_charge_dc` (+ `pv_above`), so a fix cannot be a bare `min()` — exact only for a zero-PV window. Shape to look for: hybrid inverter, `battery_rate_max_charge > inverter_limit`; `test_plan_why_reason` exercises the tooltip without asserting either way. | | An export window is in the plan but moves nothing | Fractional export limits in the sentinel band. A 99% floor combined with a low-power rung emits limits like 99.3/99.5/99.7, and both `prediction.py` and `execute.py` split at the `EXPORT_LIMIT_FREEZE`/`EXPORT_LIMIT_IDLE` sentinels, so those land in "Hold exporting" - simulator and executor agree, making it consistently inert rather than a plan/execute mismatch. With the full-power rung the same 99% floor emits exactly 99.0, the freeze sentinel itself, so a plain "export to 99%" window silently becomes a freeze (GH#4914). Related, and worth knowing before concluding an overlapping charge candidate should have won: `remove_intersecting_windows()` runs **inside every scoring simulation** (`prediction.py`), not just as a final-plan post-pass, and a freeze-export window's 99.0 limit counts as enabled and does the clipping - so the optimiser scores an overlapping charge candidate as zero benefit and "never a candidate" is accurate. The clip is shared with the kernel path (`tests/test_kernel_parity.py`'s `run_intersect_window_tests`/`run_clipping_parity_tests`), so a threshold change needs both engines updated together. | | Every ordinary export window vanishes from the plan when a saving session or Axle event is inside the horizon (battery sits full through peak PV and clips) | The export-side ratchet (GH#5221, replay-verified on main): with an event inside the horizon the automatic export threshold starts at `rate_export_min + 0.5`, and the `rate_scan_window()` rescan then **overwrites** it with `highest` — but `rate_scan_window` initialises `lowest = 99` (`fetch.py`) and only ever lowers it, so windows averaging the event price (120p) never move it and the threshold **latches at 99, not the event price** (the issue's own write-up saying "ratchets to the event price" is wrong on this point); `high_export_rates` then contains only the event windows and every ordinary export window disappears (`export_window_best = clone_windows(high_export_rates)` when `calculate_best_export` + `set_export_window` are on). **Do not read the dump's stored `rate_export_cost_threshold` as the operating threshold** — `publish_rate_and_threshold()` (`output.py`) overwrites the member from the plan's own price level (`rate_best_cost_threshold_export`), which is 120 precisely because the plan exports nowhere but the event; the June log line "High export rate found rates in range 99p to 120.0p" is the sentinel pair in the wild. Counterfactual that settles it: strip the event minutes from `rate_export` and re-run the same sequence — the threshold settles at ~20 and the PV-peak export window reappears. PR #5163 (open, for GH#5050) rewires the threshold stats to event-free and would fix this instance, but it does not touch the `lowest = 99` overwrite itself. Method note: exercise the fetch-layer threshold logic on a dump via a temporary `TEST_REGISTRY` test calling `set_rate_thresholds()` + the rescan directly (see the replay section above), not via the dump's stored values. | | Export icon/status says "off" while the battery really is exporting | `isExporting` is set at only two sites in `execute.py` (the plan's export-window branches), and the demand-freeze block appends `[Freeze exporting]` to the status text without setting it - so a demand-period freeze export publishes `binary_sensor.predbat_exporting` as "off" (GH#5125; PR #5131, open and unmerged, builds the active-vs-freeze distinction). Read the status sentence (`[Freeze exporting]`), not the binary sensor, before concluding nothing exported. | @@ -228,7 +229,7 @@ Grep for the named symbol rather than trusting a line number. - **Check a PR's base staleness before reviewing it.** A long-open PR's stored base commit can be well behind main, and a conflicting PR's headline fix may already be on main: PR #4756's `plan_iboost_smart()` fix duplicated the merged GH#4817 window-average fix, and resolving its `test_iboost.py` conflict in the PR's favour would have silently deleted main's `run_iboost_smart_average_test` — resurrecting the bug. Check `git merge-base --is-ancestor main` and diff any test-file changes against main's current tests, not the PR base, before analysing the "fix". - **GitNexus is indexed from main, so a PR branch is off-index.** `impact()` returns "Target not found" for PR-only symbols, and `detect_changes()` pins line-shifted hunks on unchanged neighbours - read its output as "which files", not "which symbols", and grep for callers instead (observed reviewing PR #5143). Same session: the review-creation API response reports `comments: 0` even when inline comments attach - verify with `gh api .../reviews//comments --jq 'length'` before resubmitting. - **`self.base.args` is process-lifetime, and Predbat writes its own dummy entity ids into it.** `self.args` is assigned once from apps.yaml (`hass.py`, `load_apps_yaml`) and never re-read for the life of the process, while `Inverter.__init__` assigns `create_entity()`'s output — `sensor.{prefix}_{inverter_type}_{id}_{entity_name}` — straight back into `self.base.args[...]`, and **those dummy ids are re-created every cycle even though the Inverter objects themselves now persist**: since PR #5126 (merged 2026-09-26, unreleased at the time of writing) `fetch_inverter_data()` builds the objects only when absent or when `num_inverters` changed, and otherwise calls `refresh_config()`, which re-reads live config and limits every cycle and still re-creates the type-named dummy entities into `self.base.args` — so anything `refresh_config()` samples from live entity state is re-sampled per cycle, not init-only, while the accumulated state `__init__` owns (the commit-once ledger `last_committed`/`commit_pending`) survives — exactly what the pre-#5126 rebuild used to wipe, which is why the committed-schedule guard never suppressed a repeated Solis/GS commit (GH#4712/#2328). **Pre-#5126 logs keep the older reading** — objects rebuilt from scratch every 5-minute cycle (`create=True` default, whole list cleared and re-created each run), so anything init-only was re-sampled per cycle (GH#5099; the pre-#5149 balance-path construction this sentence used to describe is gone — `balance_inverters()` is a pure function since PR #5149 and builds no `Inverter` objects). So with the dummy re-creation running every cycle, any "is this arg configured?" test reading `self.base.args` is reading what Predbat itself wrote last cycle — and the dummy id contains a `.`, so it passes `is_entity_id()` (PR #4745, issue #4738: a custom `GROWATTSPH` type's real `time.*` entity was overwritten by a Predbat dummy). What is lost if a guard then skips recreation on rebuild: `created_attributes` resets in `__init__` and is repopulated only by `create_entity()`, so a skipped sensor-domain write silently drops `state_class: measurement`; and because the id embeds `inverter_type`, a dummy surviving a type change names the **old** type and strands writes on a stale sensor (reproduced in a test on the PR). The obvious alternative authority, `args_from_apps_yaml`, does not generalise here: components write these same args into live `self.args` and never into the snapshot, so keying off the snapshot treats an auto-discovered inverter's select entities as unconfigured. What distinguishes them is recognising Predbat's own naming or a registry of ids created by `create_entity()` — that guard exists only on the open PR #4745, not on main. -- **Version drift.** Compare `git describe --tags` against the version in the first few lines of their log. A fair number of reports are already fixed on main, and that is a useful triage answer on its own. Two tag quirks: the `v9.0.1` and `v8.55.1` tags alias the **same commit** (`b8996659`), so "diff vs 8.55.1" is the right comparison for any v9.0.1 upgrade report from ≤v8.55.x; and PR #4991's moderate clock-skew warning postdated them both until v9.0.2 (`b8f3f956`, 2026-09-10) shipped it — a "no moderate warning" report is a pre-v9.0.2 version, not a broken feature. A third quirk (GH#5178): a `predbat.log` is **appended across add-on upgrades**, so a log attached to an issue can span two versions — one reporter's log opened with "version v8.55.0 currently running" while they were actually on v9.1.0, and a *deliberate* mid-log downgrade to A/B two versions leaves the same shape by design (GH#5335: v9.3.3 08:41 → v9.3.1 15:35 within one afternoon, giving a clean ~10-minute-apart same-scenario A/B; rebuild the timeline from the `version vX currently running` banners plus the update-service lines before attributing any behaviour to a version). Date a log by the **last** `currently running` occurrence plus restart markers (`Predbat: Startup predbat`), and cross-check `installed_version` in the debug yaml's `CONFIG_ITEMS` (`update.predbat_version` block). "GivTCP: Inverter N REST GET ... successful" lines come from the component polling REST and do not by themselves prove which Predbat version produced the surrounding lines. A fourth quirk is the **downgrade ABI crash** (GH#5213, verified twice in one log): rolling back to ≤v9.0.x with state written by v9.1.0+ crashes the first plan cycle with `TypeError: must be real number, not list` in `double_array` (`prediction_kernel.py`) — PR #5047 changed export limits from packed floats to `(mode, target, power)` tuples (kernel ABI 5→7) and 9.0.x's `double_array` cannot read them — then the next 5-minute cycle runs clean. Expect one skipped plan cycle on any downgrade path; do not misread it as the user's reported bug or as evidence the old version is broken. +- **Version drift.** Compare `git describe --tags` against the version in the first few lines of their log. A fair number of reports are already fixed on main, and that is a useful triage answer on its own. Two tag quirks: the `v9.0.1` and `v8.55.1` tags alias the **same commit** (`b8996659`), so "diff vs 8.55.1" is the right comparison for any v9.0.1 upgrade report from ≤v8.55.x; and PR #4991's moderate clock-skew warning postdated them both until v9.0.2 (`b8f3f956`, 2026-09-10) shipped it — a "no moderate warning" report is a pre-v9.0.2 version, not a broken feature. A third quirk (GH#5178): a `predbat.log` is **appended across add-on upgrades**, so a log attached to an issue can span two versions — one reporter's log opened with "version v8.55.0 currently running" while they were actually on v9.1.0, and a *deliberate* mid-log downgrade to A/B two versions leaves the same shape by design (GH#5335: v9.3.3 08:41 → v9.3.1 15:35 within one afternoon, giving a clean ~10-minute-apart same-scenario A/B; rebuild the timeline from the `version vX currently running` banners plus the update-service lines before attributing any behaviour to a version). Date a log by the **last** `currently running` occurrence plus restart markers (`Predbat: Startup predbat`), and cross-check `installed_version` in the debug yaml's `CONFIG_ITEMS` (`update.predbat_version` block). "GivTCP: Inverter N REST GET ... successful" lines come from the component polling REST and do not by themselves prove which Predbat version produced the surrounding lines. A fourth quirk is the **downgrade ABI crash** (GH#5213, verified twice in one log): rolling back to ≤v9.0.x with state written by v9.1.0+ crashes the first plan cycle with `TypeError: must be real number, not list` in `double_array` (`prediction_kernel.py`) — PR #5047 changed export limits from packed floats to `(mode, target, power)` tuples (kernel ABI 5→7) and 9.0.x's `double_array` cannot read them — then the next 5-minute cycle runs clean. Expect one skipped plan cycle on any downgrade path; do not misread it as the user's reported bug or as evidence the old version is broken. And a "worked on 9.x, broke on 9.y" claim is a version question before it is a code question — `git tag --contains` the candidate commits before accepting a regression: GH#5390's 9.3.1-works/9.3.3-broke report could not be a code regression, because the whole car-supervision family ships in v9.3.0 with identical grace constants at v9.3.1 (`git show v9.3.1:apps/predbat/const.py`) and v9.3.4's only family change ran in the user's favour. - **Log noise.** `predbat.log` carries routine `Warn:` lines (config clamps, kernel status, unsupported settings). Don't quote a warning as the root cause unless it lines up with the time the reporter describes. - **Hardware questions.** A unit test settles what the code does, not what an inverter did. Several findings here were only confirmed on live hardware. If the question is hardware behaviour, say the maintainer needs to confirm it rather than running a test to look thorough. - **`gh issue list --search` without `--repo` searches all of GitHub.** Confident-looking hits can be from unrelated projects (GH#4705). Always pass `--repo springfall2008/batpred` on a duplicate search. An **empty** result deserves the same suspicion as a confident one, and usually the cause is the query rather than the machinery: a triage run reported `gh search issues` as dead on the runner because a search for GH#4961's title text returned nothing, when that issue's real title misspells "unreasonably" - the run had silently corrected the typo while retyping the title, and searching the spelling the reporter actually used finds it (GH#4973, re-probed). Fetch the issue's exact title rather than retyping it, and sanity-check an empty result with a distinctive token from a known issue before reporting no duplicates. From 429dfe7ac8187f30d70c631d224437b75d341166 Mon Sep 17 00:00:00 2001 From: CI Date: Mon, 5 Oct 2026 00:16:40 +0100 Subject: [PATCH 2/2] docs(debug-journal): map the IOG supervision bounce symptom into the octopus row (GH#5390) Co-Authored-By: Claude Code --- tools/debug-journal.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tools/debug-journal.md b/tools/debug-journal.md index 5db3bdbf8..78fd082b4 100644 --- a/tools/debug-journal.md +++ b/tools/debug-journal.md @@ -118,7 +118,7 @@ Grep for the named symbol rather than trusting a line number. | Enphase (`enphase.py`) | Unofficial Enlighten endpoints. Accounts with MFA cannot log in at all. Discharge-to-grid schedules are required for export control. Writes need a double-submit CSRF token or return 403. Using the Enphase app at the same time can trip session limits. | `enphase_api` | | myenergi (`myenergi.py`) | `myenergi_automatic_zappi` (PR #4997) gates the Zappi half of automatic config; with it off, Zappis stay monitor-only and `control_active` is never set. `release_zappis()` has exactly one caller — `control_tick()` — reached only under `if self.control_active:` in `run()`, and `control_active` is latched at the first tick: every `enable_control()` refusal (zappi_control → automatic → automatic_zappi → enable_controls) returns *before* setting it, and the gate re-runs only when a fresh instance starts. So after a component restart with any prerequisite off, a Zappi already held in 'Stopped' stays Stopped and the `switch.*_myenergi_zappi_control` entity is not even republished (it is gated on `control_active` too) — the symptom for a future report is "car won't charge after I turned off Predbat" with a Zappi. Note `control_tick()` itself *does* release on read-only mode and on the control switch being turned off; it is the config-gate half that strands. Since PR #5150 (merged 2026-09-19) sunsynk/deye instead persist `control_active` across a restart (and re-infer it from a pre-upgrade cache), so a restart cannot strand an already-armed inverter there — the myenergi strand is the opposite direction (the gate never latching in the first place) and remains. Two adjacent lifecycle facts: published entities are **never removed** — `dashboard_item()`/`set_state()` are upsert-only (`output.py`, `ha.py` POST `/api/states`), so when a gate stops `publish_data()` publishing a switch the old entity lingers frozen at its last state until HA restarts, and there is no entity-removal path anywhere in the codebase; and `set_arg`/`set_arg_auto` writes survive a component-only restart, because `Components.restart()` stops and re-creates only the instance and never resets the shared `base.args` from apps.yaml — auto-wired values (e.g. myenergi's `car_charging_*`) persist until a full process restart. Distinguish the two restarts when checking arg lifecycle. | `myenergi` | | Gateway MQTT (`gateway.py`) | Control writes are MQTT commands the hub acknowledges. **The ack design arrived with PR #5220 and was refined by PR #5231 (both merged 2026-09-26)** — both postdate every earlier gateway log, so keep a publish-per-call reading for pre-#5220 logs: `_subscribe_acks()` subscribes to `predbat/devices//ack/+`, tolerating a broker that refuses it (a single warn via `_ack_subscribe_warned`; without the subscription or before any ack, `_send_control()` publishes exactly as before); once acks are seen, an identical command (same entity, command, payload) is published once per `_COMMAND_ACK_WINDOW` (30s), re-sent once after the unanswered window, and if that re-send is also unanswered `_acks_seen` drops so writes fall back to publish-every-call until telemetry confirms; `_process_ack()` matches any tracked id of the current value, a refusal outranks an ok (a multi-unit dispatch acks once per unit), and replay errors are handled; ids come from a clock-seeded monotonic counter (`_next_command_id()`, `PBAT`), so they never restart at `PBAT1` after a restart, and the kept-id list (`_COMMAND_ACK_IDS_KEPT`) is pruned only after the new send is out, with an ack that arrives during a failing publish still counting. **PR #5172 is a separate, still-open implementation of the same feature** (`_subscribe_command_acks`, `uuid4().hex` ids, typed outcomes threaded through service dispatch/HTTP/inverter verifier) — its symbols are not on main until it rebases, so do not start a gateway-ack triage from its names. **Two gaps around serial-less and empty status, probed on main 2026-09-25 (GH#5227, enhancement — probed with temporary tests in `test_gateway.py`, reverted after):** (1) a sole slot publishing `serial=""` (shipped firmware 1.0.0 publishes every slot, including a GivEnergy slot whose serial discovery failed) is bound as the control target: `_needs_reconfigure()` treats `""` as a newly discovered inverter, `_needs_reconfigure()` treats `""` as a newly discovered inverter and auto-config binds it — note the last-resort branch that used to fall back to `candidate_aios or list(all_inverters)` was removed in `eaef6e61` (2026-10-04), which now defers auto-config entirely with a warn when no battery-capable telemetry has arrived (`_auto_configured` stays False so `_needs_reconfigure()` retries); an empty-serial slot that does report battery telemetry is still bound, so the trap survives on that shape, and the result is empty-suffix entities (`select.predbat_gateway__charge_slot1_start`) and commands addressed to `""`, which the firmware rejects (`dongle_serial required`) — a real fleet incident on 2026-09-16 (~22 min of failed commands); an empty-serial guard in `_needs_reconfigure()`/`automatic_config()` would close it, and #5227's proposed `dongle_count` deferral does *not* cover this state (`dongle_count == len(inverters)` when the empty slot is the only one). (2) `if len(status.inverters) == 0: return` (`gateway.py`) fires **before** `_inject_entities()` (EV chargers, gateway-online) and before `_last_telemetry_time`/`update_success_timestamp()` — on a hub where all GivEnergy slots are withheld (post-predbat-gateway#335, e.g. a single-inverter hub), every status message is dropped: EV data freezes and the component fails the 60-minute staleness test (`components.py`). Symptom pointer: "gateway EV sensors froze / gateway component unhealthy while the hub is up" → this early return. Removals staying bound and reappearance triggering re-config are already the behaviour (probed). **Probe trap:** the topology branch in `automatic_config()` depends on the *whole* visible set — a probe that adds a Gateway alongside the `""` slot drops the `""` entry via the battery-presence filter (`aios` requires `battery.ByteSize() > 0`) so it never reaches the last-resort branch; construct the exact visible set the scenario implies, and note the filter silently hides data-less slots from the count in several branches (the same mechanism can make a Gateway+2-AIO fleet collapse to AIO-direct control). Existing precedent for topology probes: `TestGatewayUnitControlBinding` (`test_gateway.py`). | `gateway` | -| Octopus (`octopus.py`, `fetch.py`) | Intelligent Go tariffs are detected via `is_intelligent_go_tariff()`, and IOG-prefixed tariffs must be skipped when updating intelligent devices. Saving-session auto-join rebinding regressed when `joined_events` was empty (GH#4573). `octopus_slots_signature()` deliberately omits the time-drifting fields of active dispatch slots so a replan is not forced every cycle. `car_charging_threshold` is a strict fallback gated on `not self.car_charging_energy` (`fetch.py:228`, and `load_ml_component.py:348-361`) — it never runs as a second filter alongside a real `car_charging_energy` sensor (GH#4717). Saving-session reporting credits the full saving rate to every minute of the session on both rate tables (`load_saving_slot()`, `octopus.py:2786`); the slot dict has no baseline field to subtract (GH#2090) — a complaint that a saving session's reported total looks inflated starts here, not in a rate-fetch bug. A tariff with no `standard_unit_rates` link (IOG-TOU, GO) goes down `async_get_day_night_rates()`, which infers the off-peak window from 7 days of measurement TOU labels by plurality vote. On IOG those labels include ad-hoc dispatch slots, so a bonus slot recurring on 4+ of 7 nights became a permanent nightly cheap window and Predbat charged into it at the day rate (PR #4854 skips the inference for IOG). A "cheap slot Octopus has never heard of" report where the dispatch feeds are *empty* is this, not phantom dispatches — check the log for `Using off-peak windows [...] from measurement TOU labels`. TOU labels are UTC instants, so windows derived from them must be re-anchored for a local wall-clock tariff or they drift an hour at DST. Free-session slots have their own family of gates, separate from the saving-session ones and each found the hard way. `octopus_free_session` events used to be dropped when `code` was null, which is exactly how Octopus publishes auto-joined Weekend Happy Hours - fixed, the gate is now `start and end` and the log falls back to the event id (GH#4835). The free/saving-session args name **event** entities (`attribute="events"`/`"joined_events"`, `octopus.py`), and the BottlecapDave integration exposes Octoplus sessions as calendar + event pairs of which only the event entity carries those state attributes - the calendars have none at all, and both sides ship **disabled by default** in the integration. "The documented sensor name does not exist" plus a proposal to point at the `calendar.` twin is almost always the disabled-default trap, not a rename (GH#5370, verified against the integration's source): enable the event entity; the calendar twin cannot back these args. Matching calendar/event names differ by the `_events` suffix, so a pattern built from one domain's names mistranslates, and the deprecated pre-v17 names are scheduled for removal in January 2027. `load_free_slot()` bounded itself with `start_minutes < self.forecast_minutes` while rate minutes are indexed from `midnight_utc`, so any event more than `forecast_minutes` past midnight was dropped with no log line at all - fixed in `3bacc6e8`, the gate and the end clamp now both use `forecast_minutes + minutes_now` like `load_saving_slot()` (GH#4931). In the same function the start/end range was only updated on a successful decode while the apply block ran unconditionally, so an undecodable slot re-applied its rate over the *previous* slot's minute range (PR #4935). The join side has its own: the `joined_events` guard still ends in `saving_rate > 0` (`octopus.py:3542`), so `octopoints_per_kwh: 0` - a genuine free hour - yields no slot of any kind; probe-verified (0 gives nothing, 500 gives a saving slot, null gives nothing), and `git blame` puts that guard in `1b8c9136`, the fix for null-rate sessions (GH#3079), so it is an oversight rather than a deliberate exclusion (GH#4851). On dispatches, `rate_add_io_slots()`'s `location` check used to apply to *completed* dispatches too, and `rate_import` is rebuilt every cycle, so a retroactive AT_HOME to AWAY relabel silently un-stamped the cheap rate and `today_cost()` re-priced the whole day at the day rate (GH#4946) - **fixed in PR #4957**, which added `dispatch_billed_off_peak()` (`octopus.py:2977`): a completed dispatch is billed off-peak regardless of location, a straddling one keeps the location test, and the midday budget cap still applies. Keep the mechanism in mind when reading a log from before that merge, where the whole day reprices at the day rate. A neighbouring cap bug had no entry here at all: the daily low-rate slot budget was keyed on the *loop* minute rather than the slot start, so a window straddling noon drew a fresh 12-slot budget at 12:00 and stamped the afternoon cheap (GH#4950, **fixed in PR #4951**). The midday boundary itself is deliberate (`337db867`); only the keying was wrong. Worth knowing because "a cheap slot at an hour Octopus never offered" has at least three distinct causes in this row alone. Compare has its own rate-fetch gap: `download_octopus_rates_func()` (`octopus.py:2725`) reads only `standard-unit-rates`, and the newer IOG-SMB-FIX products are `four_rate_ev` - that endpoint answers 200 with empty results and the rates live on `day-unit-rates`/`night-unit-rates`. The main OctopusAPI component already auto-detects that product shape; compare never got the same treatment (GH#4921, probed against the live API). Lastly, `{dno_region}` in a tariff URL is substituted by `resolve_arg`'s `.format(**self.args)` (`userinterface.py:106-116`), so a missing `-` before the placeholder glues the region letter onto the product code and yields a plausible-looking 404 - check the literal URL in apps.yaml before believing a tariff has been withdrawn. Two later additions to the free/saving-session picture, one of them a correction. **eventType is carried only by the legacy `savingSessions` query path — the flexibility feed drops it (GH#4548, re-verified on main 2026-09-29 by reading `async_get_flexibility_events()`):** that function extracts only `code`/`startAt`/`endAt` from `customerFlexibilityCampaignEvents` and maps every saving event into `joinedEvents` with **no eventType**, so on the Direct path with an MPAN the joined-event loop's free-slot branch (`event_type == "WEEKEND_HAPPY_HOUR"`) can never fire and a joined Happy Hour falls through to the rewarded-saving-slot branch whenever `octopus_saving_session_rate` is set above 0. The legacy query carries eventType (the `savingSessions` GraphQL request names it and its map stores it) and is reached on the Direct path only as the flexibility path's own fallback (no MPAN, or the flexibility API returned no saving events). eventType remains the only sound discriminator between a Power Down saving session and a free Power Up/Happy Hour — the field GH#4851's `saving_rate > 0` problem needed and did not have — but *which query path served the events* decides whether you have it: "free Power Up priced as a paid saving session" on the flexibility path is first a question of provenance, and waiting for the BottlecapDave API before building on the newer feed is deliberate, not an oversight. The Direct join mutation is likewise still the legacy `joinSavingSessionsEvent`. And a genuinely counter-intuitive ordering constraint, worth reading before touching that function: Weekend Happy Hours are now skipped from `available_events` (they cannot be joined through the API - Octopus allocates them or the user books on the website), but **the skip has to sit after the reward/code/type maps are populated**, because those maps are built from the same events list and a joined Happy Hour looks its own type up there. Skip too early and the joined event has no type, takes the injected default reward, and the planner prices a free hour as an **80p/kWh saving session** (`c0fb4e9c`; the test fails if the skip is moved above the maps). Two rate-provenance reports from mid-September. **GH#5012** (car unplugged, plan still charged the house battery in the phantom cheap window): `octopus_intelligent_ignore_unplugged` correctly removed the planned dispatches, but the cheap price came from the integration's own rate sensor, not Predbat's stamping — parse the debug yaml and compare, per suspect minute, `rate_import_base` vs `rate_import_no_io` vs `io_adjusted`: a cheap price present in base and no_io with `io_adjusted` empty is the integration's rate sensor; present only after no_io is Predbat's `rate_add_io_slots()` stamping (#4516/#4950 family); `io_adjusted` populated means the minutes were flagged as adjusted — but **since PR #5304 (v9.3.3) Predbat's own `rate_add_io_slots()` flags the minutes it lowers too** (it used to write no marker at all; it mirrors the integration's `is_intelligent_adjusted`, never the fixed off-peak, time over the daily cap, or a cancelled car), so a populated `io_adjusted` no longer discriminates the integration's stamp from Predbat's own; the base-vs-`no_io` comparison still does, and `rate_replicate()` still refuses to copy `io_adjusted`-flagged minutes into future days. This attributes "integration feeds the wrong price" vs "Predbat stamps the wrong price" in one step. Nothing reverts `io_adjusted`-flagged minutes when the car is unplugged + ignore_unplugged is on — that fix would be an enhancement, blocked in practice by the integration not flagging stale adjusted entries (`io_adjusted` was empty in this dump). **`io_adjusted` itself was for a while wiped by the export/gas fetch (GH#5286, fixed in PR #5290, merged 2026-09-28):** `fetch_octopus_rates()` replaced `self.io_adjusted` on every call, so the export and gas fetches that follow the import fetch reset it to `{}` whenever `metric_octopus_export`/`metric_octopus_gas` were set — regression from #2826 dropping the old "only when adjust_key is set" guard — and with no markers left, `dynamic_load_car_strip_feed_rates()` had nothing to strip, so a car cancelled out of a dispatch still left it priced cheap and the house battery planned into it. The fix makes `self.io_adjusted` replaced only when the fetch carries an `adjust_key`; keep the mechanism for pre-#5290 logs. A companion fix (same PR) stops a compared tariff inheriting the live tariff's dispatch markers — each compared tariff now starts with none, and one left on the live import rates gets a copy. A user who instead wants the house battery limited to minutes the car actually charges has no clean knob (GH#5065): `octopus_intelligent_charging` off does **not** stop the stamping — the switch gates only car planning and vehicle prefs, planned slots are collected into `self.octopus_slots` regardless (gated only by `octopus_intelligent_ignore_unplugged`, which covers unplugged, not plugged-in-but-deferred) and are consumed unconditionally by `rate_add_io_slots()`. `octopus_slot_max: 0` is not selective either — the cap is applied in `load_octopus_slots()` as well as `rate_add_io_slots()`, so it removes the car's charging slots too. The only complete workaround is unsetting `octopus_intelligent_slot`, which loses completed-slot tracking and Octopus car planning with it. **GH#5018** (tariff switched mid-day; tomorrow's export rates stayed on the old tariff and Nordpool estimates never appeared): Predbat re-reads Octopus rate events from HA every fetch cycle (`fetch_octopus_rates`), so stale export rates are the HA integration still serving the old contract — reload the integration. Two Predbat-side traps in the same report: `futurerate` only fills *missing* minutes (the `if minute not in rates` gate in fetch.py), so stale-but-present rates always win over Nordpool estimates; and `futurerate_adjust_auto` re-runs at every FutureRate init (one per fetch cycle) and persists via `set_arg`, so mid-switch it can persist `False` over the user's manual `futurerate_adjust_export: true` — log signature `FutureRate: No futurerate adjustment enabled, skipping futurerate analysis` repeating every cycle. The debug yaml does **not** contain the Octopus rate events (grep for `day_rates` comes up empty), so replays can't reproduce integration-served rate data; use the log's `Export rates: min/max/average` lines as the provenance check (identical min/max across days = static served data, and a fixed tariff has a fixed shape while a variable one doesn't). `self.mpan` is the **import** MPAN only — `async_find_tariffs()` sets it once from the first active import agreement's `meterPoint` and never overwrites it; an account with export has import and export as separate agreements with different MPANs, and the export one is retained nowhere (PR #4972 review). The saving-session/free-electricity GraphQL queries use `self.mpan` deliberately because those campaigns are import-side; any feature that needs to identify a *meter* cannot take `self.mpan` (PR #4972, merged 2026-09-19, adds `self.tariffs[direction]["mpan"]`). Symptom: two things that should describe different supply points coming out identical — that is this, not a deduplication bug. **GH#5144** (fixed in PR #5145): the REST tariff endpoints return overlapping `DIRECT_DEBIT`/`NON_DIRECT_DEBIT` rows for the same validity window and `minute_data()` writes each row over its range, so whichever row came last in the response won — and the order is not stable across periods, so the displayed rate could flip between variants from one period to the next ("always the higher rate" was an artefact of that ordering, not a property of the bug). `filter_payment_method()` (`utils.py`) now keeps one variant at the parse point on the component, day/night, annual and minute-data paths — preferring `DIRECT_DEBIT`, keeping rows with no `payment_method` (Agile) untouched, and leaving single-variant tariffs unchanged; null `payment_method` is the common case, so any future filter must keep nulls. **The daily cheap-slot cap is derived per account but enforced per car (GH#5215, semantics unresolved):** `get_octopus_slot_max()` (`octopus.py`) takes no `car_n` — it resolves the cap from the *account's* import tariff code via `has_six_hour_cap()` (matches `IOG-SMB` → 12) or a single apps.yaml integer — while both enforcement counters are locals (`slots_per_day` in `rate_add_io_slots()` and in `load_octopus_slots()`), and both callers run once per car, so each car draws a fresh budget: up to `octopus_slot_max × num_cars` cheap half-hours priced per day. Do **not** assume the docs' per-car sentence is simply wrong: the derivation is account-level but Octopus's own blog says "*Your car gets up to 6 hours of off-peak charging per day*" (quoted in GH#4830, the multi-car confusion thread), so the code is internally inconsistent and the intended semantics is genuinely open. Exposure is invisible outside IOG-SMB — uncapped tariffs default `octopus_slot_max` to 48 — so only capped multi-car installs can see the doubling; existing tests encode the per-car behaviour, and `octopus_slot_max` is read without `index=` so per-car configuration does not exist either way. **The IOG slot-confirmation strip removes the house's cheap rate too when a car's dispatch is cancelled (GH#5335, with the #5317 family - code chain on main 57ec7bf1):** with `octopus_intelligent_dynamic` on (default), `dynamic_load_car_check()` (`plan.py`) cancels every slot of a car that sits inside a started dispatch but is not seen charging after the confirmation grace (log: `car 0 is in a dispatch but not charging, cancelling its slots`), and `dynamic_load_car_strip_feed_rates()` (`octopus.py`) re-prices the house's view of those minutes back to `rate_max_base` when the cheapness came from the integration feed, **exempting the fixed IOG band** (`OCTOPUS_NIGHT_RATE_WINDOWS["iog"]` = 23:30–05:30) — so a house battery that had planned into the dispatch loses its cheap stamp as well, and PV10 re-prices the minutes at `rate_max` outright (prediction.py's `dispatch_gone` line, `minute > 30`). Version tell: the strip is v9.3.2+ code — a reporter's deliberate mid-log downgrade left the A/B in one `predbat.log`: v9.3.3 logged `Octopus Intelligent: removed the dispatch rate from 355 minutes of cars [0] which are not charging` while v9.3.1 kept the 6.57p dispatch rate (`Charging target 2%-100%`); workaround (semantics read, not live-tested): `octopus_intelligent_dynamic` off = no confirmation and no strip, ≈9.3.1 behaviour at the cost of #5229's low-load protection. Check the car side first, though: Predbat must *see* the car charging to confirm the slot, and a `plug_status` string (`eco`/`boost`) that never equals `car_charging_now_response: charging` leaves the car "not charging" through a real charge — check the sensor's own history before blaming the planner. **#5316 is the mid-dispatch variant, fixed in PR #5319 (merged 2026-10-03) — keep the pre-#5319 signature:** a car that stops part-way through a running dispatch half hour used to have the rest of that half hour re-priced at the day rate; the strip paths (`rate_add_io_slots()`, `dynamic_load_car_strip_feed_rates()`) now start from `dynamic_load_car_strip_from()` (`plan.py`) — the end of the half hour the car was last seen charging in, which is billed off-peak in full by Octopus — and the confirmation is never cleared, only outlived. Review-round refinements ride along: a sensor reading confirms a half hour only from 2 minutes into it, and the grace clock is not kept across a restart (saved state is judged on the skewed clock). | `octopus_*`, `saving_session*` | +| Octopus (`octopus.py`, `fetch.py`) | Intelligent Go tariffs are detected via `is_intelligent_go_tariff()`, and IOG-prefixed tariffs must be skipped when updating intelligent devices. Saving-session auto-join rebinding regressed when `joined_events` was empty (GH#4573). `octopus_slots_signature()` deliberately omits the time-drifting fields of active dispatch slots so a replan is not forced every cycle. `car_charging_threshold` is a strict fallback gated on `not self.car_charging_energy` (`fetch.py:228`, and `load_ml_component.py:348-361`) — it never runs as a second filter alongside a real `car_charging_energy` sensor (GH#4717). Saving-session reporting credits the full saving rate to every minute of the session on both rate tables (`load_saving_slot()`, `octopus.py:2786`); the slot dict has no baseline field to subtract (GH#2090) — a complaint that a saving session's reported total looks inflated starts here, not in a rate-fetch bug. A tariff with no `standard_unit_rates` link (IOG-TOU, GO) goes down `async_get_day_night_rates()`, which infers the off-peak window from 7 days of measurement TOU labels by plurality vote. On IOG those labels include ad-hoc dispatch slots, so a bonus slot recurring on 4+ of 7 nights became a permanent nightly cheap window and Predbat charged into it at the day rate (PR #4854 skips the inference for IOG). A "cheap slot Octopus has never heard of" report where the dispatch feeds are *empty* is this, not phantom dispatches — check the log for `Using off-peak windows [...] from measurement TOU labels`. TOU labels are UTC instants, so windows derived from them must be re-anchored for a local wall-clock tariff or they drift an hour at DST. Free-session slots have their own family of gates, separate from the saving-session ones and each found the hard way. `octopus_free_session` events used to be dropped when `code` was null, which is exactly how Octopus publishes auto-joined Weekend Happy Hours - fixed, the gate is now `start and end` and the log falls back to the event id (GH#4835). The free/saving-session args name **event** entities (`attribute="events"`/`"joined_events"`, `octopus.py`), and the BottlecapDave integration exposes Octoplus sessions as calendar + event pairs of which only the event entity carries those state attributes - the calendars have none at all, and both sides ship **disabled by default** in the integration. "The documented sensor name does not exist" plus a proposal to point at the `calendar.` twin is almost always the disabled-default trap, not a rename (GH#5370, verified against the integration's source): enable the event entity; the calendar twin cannot back these args. Matching calendar/event names differ by the `_events` suffix, so a pattern built from one domain's names mistranslates, and the deprecated pre-v17 names are scheduled for removal in January 2027. `load_free_slot()` bounded itself with `start_minutes < self.forecast_minutes` while rate minutes are indexed from `midnight_utc`, so any event more than `forecast_minutes` past midnight was dropped with no log line at all - fixed in `3bacc6e8`, the gate and the end clamp now both use `forecast_minutes + minutes_now` like `load_saving_slot()` (GH#4931). In the same function the start/end range was only updated on a successful decode while the apply block ran unconditionally, so an undecodable slot re-applied its rate over the *previous* slot's minute range (PR #4935). The join side has its own: the `joined_events` guard still ends in `saving_rate > 0` (`octopus.py:3542`), so `octopoints_per_kwh: 0` - a genuine free hour - yields no slot of any kind; probe-verified (0 gives nothing, 500 gives a saving slot, null gives nothing), and `git blame` puts that guard in `1b8c9136`, the fix for null-rate sessions (GH#3079), so it is an oversight rather than a deliberate exclusion (GH#4851). On dispatches, `rate_add_io_slots()`'s `location` check used to apply to *completed* dispatches too, and `rate_import` is rebuilt every cycle, so a retroactive AT_HOME to AWAY relabel silently un-stamped the cheap rate and `today_cost()` re-priced the whole day at the day rate (GH#4946) - **fixed in PR #4957**, which added `dispatch_billed_off_peak()` (`octopus.py:2977`): a completed dispatch is billed off-peak regardless of location, a straddling one keeps the location test, and the midday budget cap still applies. Keep the mechanism in mind when reading a log from before that merge, where the whole day reprices at the day rate. A neighbouring cap bug had no entry here at all: the daily low-rate slot budget was keyed on the *loop* minute rather than the slot start, so a window straddling noon drew a fresh 12-slot budget at 12:00 and stamped the afternoon cheap (GH#4950, **fixed in PR #4951**). The midday boundary itself is deliberate (`337db867`); only the keying was wrong. Worth knowing because "a cheap slot at an hour Octopus never offered" has at least three distinct causes in this row alone. Compare has its own rate-fetch gap: `download_octopus_rates_func()` (`octopus.py:2725`) reads only `standard-unit-rates`, and the newer IOG-SMB-FIX products are `four_rate_ev` - that endpoint answers 200 with empty results and the rates live on `day-unit-rates`/`night-unit-rates`. The main OctopusAPI component already auto-detects that product shape; compare never got the same treatment (GH#4921, probed against the live API). Lastly, `{dno_region}` in a tariff URL is substituted by `resolve_arg`'s `.format(**self.args)` (`userinterface.py:106-116`), so a missing `-` before the placeholder glues the region letter onto the product code and yields a plausible-looking 404 - check the literal URL in apps.yaml before believing a tariff has been withdrawn. Two later additions to the free/saving-session picture, one of them a correction. **eventType is carried only by the legacy `savingSessions` query path — the flexibility feed drops it (GH#4548, re-verified on main 2026-09-29 by reading `async_get_flexibility_events()`):** that function extracts only `code`/`startAt`/`endAt` from `customerFlexibilityCampaignEvents` and maps every saving event into `joinedEvents` with **no eventType**, so on the Direct path with an MPAN the joined-event loop's free-slot branch (`event_type == "WEEKEND_HAPPY_HOUR"`) can never fire and a joined Happy Hour falls through to the rewarded-saving-slot branch whenever `octopus_saving_session_rate` is set above 0. The legacy query carries eventType (the `savingSessions` GraphQL request names it and its map stores it) and is reached on the Direct path only as the flexibility path's own fallback (no MPAN, or the flexibility API returned no saving events). eventType remains the only sound discriminator between a Power Down saving session and a free Power Up/Happy Hour — the field GH#4851's `saving_rate > 0` problem needed and did not have — but *which query path served the events* decides whether you have it: "free Power Up priced as a paid saving session" on the flexibility path is first a question of provenance, and waiting for the BottlecapDave API before building on the newer feed is deliberate, not an oversight. The Direct join mutation is likewise still the legacy `joinSavingSessionsEvent`. And a genuinely counter-intuitive ordering constraint, worth reading before touching that function: Weekend Happy Hours are now skipped from `available_events` (they cannot be joined through the API - Octopus allocates them or the user books on the website), but **the skip has to sit after the reward/code/type maps are populated**, because those maps are built from the same events list and a joined Happy Hour looks its own type up there. Skip too early and the joined event has no type, takes the injected default reward, and the planner prices a free hour as an **80p/kWh saving session** (`c0fb4e9c`; the test fails if the skip is moved above the maps). Two rate-provenance reports from mid-September. **GH#5012** (car unplugged, plan still charged the house battery in the phantom cheap window): `octopus_intelligent_ignore_unplugged` correctly removed the planned dispatches, but the cheap price came from the integration's own rate sensor, not Predbat's stamping — parse the debug yaml and compare, per suspect minute, `rate_import_base` vs `rate_import_no_io` vs `io_adjusted`: a cheap price present in base and no_io with `io_adjusted` empty is the integration's rate sensor; present only after no_io is Predbat's `rate_add_io_slots()` stamping (#4516/#4950 family); `io_adjusted` populated means the minutes were flagged as adjusted — but **since PR #5304 (v9.3.3) Predbat's own `rate_add_io_slots()` flags the minutes it lowers too** (it used to write no marker at all; it mirrors the integration's `is_intelligent_adjusted`, never the fixed off-peak, time over the daily cap, or a cancelled car), so a populated `io_adjusted` no longer discriminates the integration's stamp from Predbat's own; the base-vs-`no_io` comparison still does, and `rate_replicate()` still refuses to copy `io_adjusted`-flagged minutes into future days. This attributes "integration feeds the wrong price" vs "Predbat stamps the wrong price" in one step. Nothing reverts `io_adjusted`-flagged minutes when the car is unplugged + ignore_unplugged is on — that fix would be an enhancement, blocked in practice by the integration not flagging stale adjusted entries (`io_adjusted` was empty in this dump). **`io_adjusted` itself was for a while wiped by the export/gas fetch (GH#5286, fixed in PR #5290, merged 2026-09-28):** `fetch_octopus_rates()` replaced `self.io_adjusted` on every call, so the export and gas fetches that follow the import fetch reset it to `{}` whenever `metric_octopus_export`/`metric_octopus_gas` were set — regression from #2826 dropping the old "only when adjust_key is set" guard — and with no markers left, `dynamic_load_car_strip_feed_rates()` had nothing to strip, so a car cancelled out of a dispatch still left it priced cheap and the house battery planned into it. The fix makes `self.io_adjusted` replaced only when the fetch carries an `adjust_key`; keep the mechanism for pre-#5290 logs. A companion fix (same PR) stops a compared tariff inheriting the live tariff's dispatch markers — each compared tariff now starts with none, and one left on the live import rates gets a copy. A user who instead wants the house battery limited to minutes the car actually charges has no clean knob (GH#5065): `octopus_intelligent_charging` off does **not** stop the stamping — the switch gates only car planning and vehicle prefs, planned slots are collected into `self.octopus_slots` regardless (gated only by `octopus_intelligent_ignore_unplugged`, which covers unplugged, not plugged-in-but-deferred) and are consumed unconditionally by `rate_add_io_slots()`. `octopus_slot_max: 0` is not selective either — the cap is applied in `load_octopus_slots()` as well as `rate_add_io_slots()`, so it removes the car's charging slots too. The only complete workaround is unsetting `octopus_intelligent_slot`, which loses completed-slot tracking and Octopus car planning with it. **GH#5018** (tariff switched mid-day; tomorrow's export rates stayed on the old tariff and Nordpool estimates never appeared): Predbat re-reads Octopus rate events from HA every fetch cycle (`fetch_octopus_rates`), so stale export rates are the HA integration still serving the old contract — reload the integration. Two Predbat-side traps in the same report: `futurerate` only fills *missing* minutes (the `if minute not in rates` gate in fetch.py), so stale-but-present rates always win over Nordpool estimates; and `futurerate_adjust_auto` re-runs at every FutureRate init (one per fetch cycle) and persists via `set_arg`, so mid-switch it can persist `False` over the user's manual `futurerate_adjust_export: true` — log signature `FutureRate: No futurerate adjustment enabled, skipping futurerate analysis` repeating every cycle. The debug yaml does **not** contain the Octopus rate events (grep for `day_rates` comes up empty), so replays can't reproduce integration-served rate data; use the log's `Export rates: min/max/average` lines as the provenance check (identical min/max across days = static served data, and a fixed tariff has a fixed shape while a variable one doesn't). `self.mpan` is the **import** MPAN only — `async_find_tariffs()` sets it once from the first active import agreement's `meterPoint` and never overwrites it; an account with export has import and export as separate agreements with different MPANs, and the export one is retained nowhere (PR #4972 review). The saving-session/free-electricity GraphQL queries use `self.mpan` deliberately because those campaigns are import-side; any feature that needs to identify a *meter* cannot take `self.mpan` (PR #4972, merged 2026-09-19, adds `self.tariffs[direction]["mpan"]`). Symptom: two things that should describe different supply points coming out identical — that is this, not a deduplication bug. **GH#5144** (fixed in PR #5145): the REST tariff endpoints return overlapping `DIRECT_DEBIT`/`NON_DIRECT_DEBIT` rows for the same validity window and `minute_data()` writes each row over its range, so whichever row came last in the response won — and the order is not stable across periods, so the displayed rate could flip between variants from one period to the next ("always the higher rate" was an artefact of that ordering, not a property of the bug). `filter_payment_method()` (`utils.py`) now keeps one variant at the parse point on the component, day/night, annual and minute-data paths — preferring `DIRECT_DEBIT`, keeping rows with no `payment_method` (Agile) untouched, and leaving single-variant tariffs unchanged; null `payment_method` is the common case, so any future filter must keep nulls. **The daily cheap-slot cap is derived per account but enforced per car (GH#5215, semantics unresolved):** `get_octopus_slot_max()` (`octopus.py`) takes no `car_n` — it resolves the cap from the *account's* import tariff code via `has_six_hour_cap()` (matches `IOG-SMB` → 12) or a single apps.yaml integer — while both enforcement counters are locals (`slots_per_day` in `rate_add_io_slots()` and in `load_octopus_slots()`), and both callers run once per car, so each car draws a fresh budget: up to `octopus_slot_max × num_cars` cheap half-hours priced per day. Do **not** assume the docs' per-car sentence is simply wrong: the derivation is account-level but Octopus's own blog says "*Your car gets up to 6 hours of off-peak charging per day*" (quoted in GH#4830, the multi-car confusion thread), so the code is internally inconsistent and the intended semantics is genuinely open. Exposure is invisible outside IOG-SMB — uncapped tariffs default `octopus_slot_max` to 48 — so only capped multi-car installs can see the doubling; existing tests encode the per-car behaviour, and `octopus_slot_max` is read without `index=` so per-car configuration does not exist either way. **The IOG slot-confirmation strip removes the house's cheap rate too when a car's dispatch is cancelled (GH#5335, with the #5317 family - code chain on main 57ec7bf1):** with `octopus_intelligent_dynamic` on (default), `dynamic_load_car_check()` (`plan.py`) cancels every slot of a car that sits inside a started dispatch but is not seen charging after the confirmation grace (log: `car 0 is in a dispatch but not charging, cancelling its slots`), and `dynamic_load_car_strip_feed_rates()` (`octopus.py`) re-prices the house's view of those minutes back to `rate_max_base` when the cheapness came from the integration feed, **exempting the fixed IOG band** (`OCTOPUS_NIGHT_RATE_WINDOWS["iog"]` = 23:30–05:30) — so a house battery that had planned into the dispatch loses its cheap stamp as well, and PV10 re-prices the minutes at `rate_max` outright (prediction.py's `dispatch_gone` line, `minute > 30`). Version tell: the strip is v9.3.2+ code — a reporter's deliberate mid-log downgrade left the A/B in one `predbat.log`: v9.3.3 logged `Octopus Intelligent: removed the dispatch rate from 355 minutes of cars [0] which are not charging` while v9.3.1 kept the 6.57p dispatch rate (`Charging target 2%-100%`); workaround (semantics read, not live-tested): `octopus_intelligent_dynamic` off = no confirmation and no strip, ≈9.3.1 behaviour at the cost of #5229's low-load protection. Check the car side first, though: Predbat must *see* the car charging to confirm the slot, and a `plug_status` string (`eco`/`boost`) that never equals `car_charging_now_response: charging` leaves the car "not charging" through a real charge — check the sensor's own history before blaming the planner. **#5316 is the mid-dispatch variant, fixed in PR #5319 (merged 2026-10-03) — keep the pre-#5319 signature:** a car that stops part-way through a running dispatch half hour used to have the rest of that half hour re-priced at the day rate; the strip paths (`rate_add_io_slots()`, `dynamic_load_car_strip_feed_rates()`) now start from `dynamic_load_car_strip_from()` (`plan.py`) — the end of the half hour the car was last seen charging in, which is billed off-peak in full by Octopus — and the confirmation is never cleared, only outlived. Review-round refinements ride along: a sensor reading confirms a half hour only from 2 minutes into it, and the grace clock is not kept across a restart (saved state is judged on the skewed clock). The bounce symptom itself — the battery flipping Charge↔Demand while a car charges in bursts — is the supervision working as coded (GH#5390, log-verified): a non-"charging" reading counts after `DYNAMIC_LOAD_CAR_START_MINUTES` (3 min) inside the dispatch, the cheap dispatch rate is stripped from the car's still-future slots (house plan re-prices those minutes) and it resumes on the next "charging" reading — ground-truth with the hourly `Last hour: ... car X kWh` log lines against Octopus's own per-slot amounts before calling it a bug. | `octopus_*`, `saving_session*` | | Kraken / EDF (`kraken.py`) | EDF Kraken answers a day/night-structured tariff with **HTTP 400** on `standard-unit-rates/` (`{"detail": "This tariff has day and night rates, not standard."}`) while `day-unit-rates`/`night-unit-rates` return 200 and standing charges 200 — where Octopus four-rate products (GH#4921) answer the same-shaped URL **200 with empty results**. So the "empty results, then data empty" signature belongs to `octopus.py` and a hard 400 to `kraken.py`; do not carry the expectation across (GH#5166, probed against the live API). A nonexistent product still 404s, so 404 and 400 are both live "REST can't serve this tariff" signals on EDF — probe the exact URL from the reporter's log before assuming either. **Fixed in PR #5167** (`async_fetch_rates()`): the GraphQL `applicableRates` fallback now gates on `KRAKEN_REST_RATES_UNAVAILABLE_STATUSES` = (400, 404, 410), per direction (404 covers both a private product and one retired from the REST API; the authenticated retry stays 404/410 because auth cannot change a 400), and only counts a failure when the fallback comes back empty without counting one — `async_graphql_query()` counts its own failures, so a caller that also counts scores 2 per down cycle; snapshot the counter around the call. The standing-charge fallback stays deliberately 404/410-only (`kraken.py` — "400 is a rates-endpoint-only answer"). Two traps from the fix's review: `_fetch_rates_rest()` returns `(None, None)` on a network error, so `err is None` does **not** mean success — success is `(results, None)`; the merged code stashes the public attempt's status and restores it when the authenticated retry hits a network error (pre-fix symptom: a private-product tariff's rates vanish for one cycle with no fallback log line, while an identical 404 one cycle later recovers). The day/night endpoints' rows carry the same `payment_method` variant overlap as GH#5144 — suspected, no probe; check `minute_data`'s handling before consuming them directly. Also benign: `Warn: Kraken: Auth not available for find-tariffs` repeating on early restarts before the first `Tariff discovered` line is setup-phase, not a second bug. | `kraken` | | Axle (`axle.py`) | Export sessions have to boost the import rate as well as the export rate - introduced deliberately by PR #4520 (first release v8.48.2, mirrored on Octopus saving sessions), documented in docs/energy-rates.md, and upheld by the maintainer on #5060. Two reports of the same design now exist (#5060 closed as dup-of-#4277 with the by-design answer given; #5175 triaged as enhancement, no regression - `tools/triage_test.sh axle` asserts the dual boost), and **no config knob disables just the import side** (`axle_pence_per_kwh` scales both directions; config.py has only `axle_api_key`/`axle_pence_per_kwh`/`axle_automatic`/`axle_control`). The "textual plans / colour coding skewed" half of such reports belongs to the #5050 threshold row (boosts before `rate_scan()`), not here. State is published unconditionally from `run()` so a fetch failure does not freeze the sensor at a stale value. `load_axle_slot()` has no lower time bound - it checks only `start_minutes < forecast_minutes + minutes_now`, not the `start_minutes >= 0` guard `load_free_slot()` has (octopus.py) - so a closed Axle event writes dead keys at negative minutes through `rate_dict.get(minute, 0)` (GH#5036, replay-verified against the reporter's dump: the +100 boosts sit only at the two closed events' actual UTC start/end times). The dead keys are inert in the plan - `rate_minmax()` and the window scans start at `minutes_now` - so "a dead event projected forward as free export" is not what the code does. The comment at octopus.py:2859 saying `load_saving_slot()` and `load_axle_slot()` "both bound themselves that way" is half-wrong: axle bounds only its end, and `load_saving_slot()` also lacks the start guard but is harmless there because its inner loop writes only `if minute in rate_dict`. Two September reports extend this row. **GH#5060** (closed as a duplicate of #4277, the same defect reported a year earlier and never root-caused — #4277 covers Octopus saving sessions too): event boosts are applied *before* user rate overrides on both sides — `load_axle_slot(..., export=True)` runs ahead of `basic_rates(rates_export_override, ...)` (import mirrors it: `load_saving_slot`/`load_axle_slot` then the import override) — so a full-horizon `rates_export_override` rewrites the event minutes and re-marks them `user`. That reporter had no `rates_export` key in apps.yaml at all, so the override was their only export source, which is why the import-side boost survived while the export-side +100 did not. Workaround (maintainer's suggestion on #4277, checked viable against the code): move the fixed schedule from `rates_export_override` into `rates_export` — it becomes the base tariff via `basic_rates` before the boosts run, so the event stacks on top and nothing downstream rewrites it (assuming manual export rates are empty, since `apply_manual_rates` also runs after the boosts). Forensic technique: in a debug yaml, the `rate_*_replicated` dict discriminates the two failure modes in one step — mark `user` through the event window ⇒ the boost ran and a later override clobbered it; the `saving` marks missing while the boost is missing ⇒ the boost call never ran (or, per GH#5036 above, wrote dead keys). Compare against `rate_*_base` for the tariff's own price; same idea as the GH#5012 base/no_io/adjusted comparison in the Octopus row. **GH#5050** (shared with the Octopus row): the same boosts run before `rate_scan()`, so the automatic low-rate threshold classifies the whole day low — see the symptom-table row on `set_rate_thresholds()`. | `axle` | | History fetch / memory (`ha.py`) | History is fetched in `HISTORY_CHUNK_DAYS`-sized chunks with boundary dedup — records landing exactly on a chunk start inside a data gap corrupted smoothing before that was fixed. The largest memory peak in a run is ML load-predictor training (`load_predictor.py`), not the plan. The same chunking is the REST-call amplifier that can trip the fatal API-error counter at startup — see the "Too many API errors" symptom row. | `history_chunking` |