Skip to content

docs(debug-journal): fold in the 2026-10-01/02 queue slice (5 candidates); mark GH#5068/#5136-integer/#5222/#5304-branch fixed, split GH#5136's live publish_info half, correct the GH#5012 io_adjusted heuristic - #5342

Merged
springfall2008 merged 1 commit into
mainfrom
bot/debug-journal-2026-10-02
Oct 2, 2026

Conversation

@springfall2008

Copy link
Copy Markdown
Owner

Automated PR — the daily journal-update flow, opening its own PR so the daemon's copy on bot/debug-journal-2026-10-02 gets real content (run log in the workflow that pushed this branch). Base: main @ 57ec7bf (v9.3.3).

All 5 queued candidates were in scope (limit 15); 0 left queued. Every candidate and every correction below was verified against current main 57ec7bf (also the v9.3.3 tag) before folding in. 5/5 candidates folded in, none dropped as pure duplicates; all five also updated the journal where a fix had merged since the finding was written.

Folded in

  • GH#4572 — "Hold for car" is expected; no native EV-SoC or price gate exists (folded in). Car-charging row: the hold only stops the battery feeding the car — the car's grid charging and the charger's start/stop are outside its scope, so "battery idle while the EV charges on PV-then-grid" is the intended steady state; the full car config surface has no EV-SoC gate and no enforceable price threshold (car_charging_soc/_limit are input sensors, not thresholds), "charge the EV only when X" is an enhancement ask; in-tree tooling direction = open PR Expose solar surplus power and force-export window state as sensors #3791 (predbat.solar_surplus_power + binary_sensor.predbat_force_export_slot, sensor-only pared-down form, still unmerged — owner review round outstanding 2026-10-02). Also a new symptom-table row ("Battery sits idle ('Hold for car') while the car charges from the grid"), so the expected-behaviour answer is findable from the symptom side.
  • GH#5330 — house-battery charge placement inside a long IOG dispatch is a cost tie (folded in). New symptom-table row next to the fix(octopus): model a vanishing IOG dispatch for octopus_intelligent_slot users #5304 row: within-block placement is metric-neutral (probe-verified on the reporter's own replay: 148.5602p / 148.5586p / 148.5745p across charged-early / continuous / the reporter's late ramp), the IOG earlier-charge skew (IO_ADJUST_* constants, _io_rate_adjustment() inside sort_window_by_price_combined()) is ranking-level only, reordering candidate windows but never a metric tie-break; the probe harness (replay + direct run_prediction() variants, since removed from the tree) is kept as the A/B technique; PR fix(octopus): model a vanishing IOG dispatch for octopus_intelligent_slot users #5304's PV10 hedge steering placement is kept explicitly suspected, not verified.
  • GH#5334 — GivTCP REST: a missing inverter-details block silently stales published sensors and fakes a clock-skew alarm (folded in). GivTCP REST row: inverter_details() returns {} with no log (normal on v3) when the gateway omits raw.invertor.serial_number; publish_data()'s if value: gates freeze each detail-block sensor at its last state (stale identical to live) while serial_number is published unconditionally — the serial sensor reading unknown while inverter_time stays frozen is the tell; check_clock_skew() then crosses the ≥30-min restart threshold every cycle ("Clock skew >= 30 minutes command []" with the clock actually fine). Upstream: Give the Gateway and EMS a serial number in the REST raw data britkat1980/giv_tcp#597, unmerged at triage time. Discussed fix shapes + the generalisation: any "Predbat-published entity looks frozen but claims to be live" report on a REST-backed component is this publish-gate class.
  • GH#5335 — IOG slot-confirmation strip + the foreign debug yaml, split across three entries (folded in).
    • Octopus row: the slot-confirmation strip chain — dynamic_load_car_check() cancels all slots of a car in a started dispatch not seen charging after the grace (car 0 is in a dispatch but not charging, cancelling its slots), dynamic_load_car_strip_feed_rates() re-prices the house's view back to rate_max_base (integration-feed cheapness, fixed IOG band 23:30–05:30 exempt), PV10 re-prices io_adjusted minutes >30 min ahead at rate_max; log A/B from the reporter's mid-log downgrade (v9.3.3 strips, v9.3.1 kept the 6.57p dispatch rate); workaround octopus_intelligent_dynamic off; the car-side confound — a plug_status string (eco/boost) that never matches car_charging_now_response: charging keeps the car "not charging" through a real charge, check the sensor's history before blaming the planner.
    • New traps bullet: cross-check an attached debug yaml against the attached log before believing either — SoC/soc_max and resolved entity serials at the yaml's capture minute; a foreign yaml (re-shared from a community thread) can carry different SoC, Zappi/GE serials and Octopus accounts; only the log is then usable and the reporter's own yaml must be re-asked for. battery exporting while EV charging and battery set to demand in cheap rate with IOG #5317 was a prior sighting (prior session folded a partial note; this is the durable rule).
    • Version-drift trap: a deliberate mid-log downgrade to A/B two versions leaves both versions' behaviour in one log — rebuild the timeline from the version vX currently running banners + update-service lines (GH#5335's clean ~10-min-apart A/B).
  • GH#5336 — Load ML's hardcoded 48h horizon leaves the plan's load at zero beyond it (folded in). The Load ML row (renamed from "Load ML CPU spikes" since it now holds non-CPU findings): PREDICT_HORIZON (load_predictor.py:35) feeds only predict()/logs/save-metadata — training is independent, load() validates only version + architecture, so models survive a horizon change; with load_ml_source on, load_forecast_only skips the weighted-bucket forecast wholesale so the plan's load past minute 2880 is exactly 0 (log tell: Starting autoregressive prediction loop for 576 steps (48.0 hours)); the "over 48h" in the Generated N predictions log string is hardcoded in the string, not the model; test_load_ml.py pins 576 steps; enhancement (configurable 48/72/96), no fallback beyond the ML horizon by design of load_forecast_only.

Corrections to existing entries (overtaken by merges since the last flush)

Verification & gate

  • Statuses re-checked this session: PR Run ML training off the event loop and let a shutdown abandon it #5112 (load-ML shutdown) and Expose solar surplus power and force-export window state as sensors #3791 (solar-surplus sensors) both still open — the entries citing them stand.
  • Pre-commit hooks: 14/14 green (cspell needed "britkat1980" added to .cspell/custom-dictionary-workspace.txt — repo names have precedent in the dict, e.g. "hultenvp"). The quick suite inside run_pre_commit aborted at 00:14 BST on test_teslemetry_local_weekday_follows_the_base_clock's assert (datetime.now().weekday() vs the UTC fallback) — exactly the midnight-window failure the journal's own trap bullet documents (00:09/00:12 BST sightings on clean trees); unrelated to this docs-only diff, and the unhandled raise also terminates the later tests (also documented). The docs-relevant hooks all pass; a maintainer rerunning ./run_all --quick outside the 00:00–01:00 BST window should see the suite green.

…tes); mark GH#5068/#5136-integer/#5222/#5304-branch fixed, split GH#5136's live publish_info half, correct the GH#5012 io_adjusted heuristic

Co-Authored-By: Claude Code <noreply@anthropic.com>
@springfall2008
springfall2008 marked this pull request as ready for review October 2, 2026 07:15
Copilot AI balanced review requested due to automatic review settings October 2, 2026 07:15
@springfall2008
springfall2008 merged commit 549e476 into main Oct 2, 2026
2 checks passed
@springfall2008
springfall2008 deleted the bot/debug-journal-2026-10-02 branch October 2, 2026 07:15

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Three journal entries contain inaccurate or misleading current-behaviour documentation.

Review effort: Balanced
Findings: 3 Low severity

Open (3)
What changed in this PR

Updates the debugging journal with recent triage findings and merged-fix status.

Changes:

  • Adds five investigation findings and symptom guidance.
  • Updates several entries for recently merged fixes.
  • Adds the upstream username to the spelling dictionary.
File Description
tools/​debug-journal.md Updates diagnostic findings and fix statuses.
.cspell/​custom-dictionary-workspace.txt Adds britkat1980.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread tools/debug-journal.md
@@ -103,22 +103,22 @@ Grep for the named symbol rather than trusting a line number.
| Area | What past debugging found | Targeted test |
|------|---------------------------|---------------|
| Fox (`fox.py`) | The cloud API returns errno 42015/44096 for settings a given device does not support (`FOX_SETTINGS_UNSUPPORTED_ERRNO`); those are marked unavailable and never polled or written again. Entity type matters — WorkMode is a select, ExportLimit a number. Two later capacity/schedule traps, both still live: `publish_data()` sums every `batteryList` entry unconditionally (`fox.py:1779-1781`), and an AIO ESS returns one physical pack as four `bmu` entries all carrying the inverter's own `batterySN`, so `soc_max` comes out at 4x the correct `batteryDesignCapacity` and `soc_kw` with it (GH#4919, read out of the reporter's own API response). `fox_automatic: true` re-`set_arg`s `soc_max` every cycle and `soc_max` is not in `CONFIG_API_OVERRIDE`, so apps.yaml cannot override it - the escape hatch is `fox_automatic: false`. With the Mode Scheduler deleted the scheduler `enable` flag reads 0 and `compute_schedule()` (`fox.py:1080`) derives the displayed charge window from the legacy `forceChargeTime` read, which Predbat cannot clear: `set_battery_charging_time()` (`fox.py:1009`), the only writer of that endpoint, has zero production callers - re-verified on main, where it is referenced only by `test_fox_api.py` (GH#4939). Note this is the *residual* half of that issue: the blocker the reporter actually hit was HA event routing filtering on the literal string `predbat` rather than the configured entity prefix, **fixed in PR #4962**. Unlike the self-referential HA verify trap below, a Fox `didn't complete got X` is a genuine failure - the read-back polls entities the component republishes from Fox Cloud. A separate trap on the **HA/modbus path** (inverter type `FoxESS`, not `FoxCloud`): the stock template binds `reserve:` to `number.foxess_min_soc_on_grid` (`templates/fox.yaml`), and that file's own mode table ties both Freeze charging and Hold charging to that entity. A reporter had `reserve:` bound to `number.foxess_inverter_min_soc` instead - a name that appears nowhere in `templates/` or `docs/` - so Predbat never read or wrote `min_soc_on_grid`; when the inverter stranded it at 100 (observed sitting there six days) Predbat could neither see it nor clear it, and `min_soc_on_grid` at 100% stops the battery discharging while grid-connected. `FoxESS` is `has_idle_time: False` (`config.py`), so the `idle time is ...` line is bookkeeping only and never programs a demand period - discharge outside a charge window depends entirely on the inverter's own settings, which is what makes a wrong reserve binding silent (GH#4961). FoxCloud freeze export (GH#5015, **fixed in PR #5038**, v9.0.2 - keep the mechanism for pre-v9.0.2 logs): `support_feedin_first: True` makes `prediction.py`'s freeze branch model freeze export as genuine Feed-in-First — load exports up to the limit, only the surplus beyond it charges the battery — but the cloud path used to never select the `Feedin` work mode: `apply_battery_schedule()` (`fox.py`) only emitted `SelfUse`/`ForceCharge`/`ForceDischarge` groups and `adjust_inverter_mode()` hardcodes `SelfUse` for Fox (`inverter.py`), so the inverter sat in SelfUse where surplus PV charges the battery *before* exporting and the plan's freeze-export revenue didn't match the hardware. #5038 makes `apply_battery_schedule()` select `Feedin` as the baseline work mode whenever freeze export is requested; note the cloud API's verify read is slow there — a write can log success, read back pre-write 3s later and read back the new mode after ~18s (recorded in a `fox.py` comment), so a quick verify read is not evidence of failure. `support_feedin_first` therefore means two different things depending on connection method: the modelling is correct on the FoxESS HA/modbus path, where `discharge_freeze_service` genuinely selects Feed-in First (#4207, fixed by #4425), while #4582 extended the flag to the cloud defs with no equivalent execution path. Log check: a freeze-export day shows hundreds of full-day `SelfUse` groups in `Fox: New schedule` lines and never a `ForceDischarge` or `Feedin` group — and `ForceDischarge` appearing in a `Fetch scheduler V1 returned` enum list is **not** evidence a discharge window was written. Workaround: turn `set_export_freeze` off so the planner stops selecting freeze-export slots. #4182 is the neighbouring open issue. | `fox_api`, `fox_oauth` |
| Solis (`solis.py`) | `SOLIS_CID_STORAGE_MODE = 636` is Modbus 43110, a bit mask. On firmware "4B and above" (what `is_tou_v2_mode()` detects: CID 6798 reads 43605) the timed charge/discharge enable moved to the per-slot registers — `SOLIS_CID_CHARGE_ENABLE_BASE`/`..._DISCHARGE_ENABLE_BASE`, Modbus 43707, six slots not three — and every mode value carrying TOU bit 1 (3/35/43/51/98) was dropped; 35 became 33, 98 became 96. Such an inverter answers a CID 636 write with code 0 and reads back without bit 1, so a log full of CID 636 verification warnings is a refused bit, not a failed write (GH#4707). `set_storage_mode_if_needed()` decides the bit from `is_tou_v2_mode()`: never asked for on V2, always asked for and retried on V1, where bit 1 still *is* the timed charge/discharge enable. Do not try to learn it from the read-back instead — GH#4710 did, and GH#4774 showed why that cannot work: on six inverters the same write of 179 verifies minutes before and minutes after the one that reads back 177, so a single post-write read is not evidence of a firmware property. Latching it was also silent rather than noisy: the verify read refreshes `cached_values`, so after a stripped write the cache held the stripped value, the suppressed computed value matched it, and `set_storage_mode_if_needed()` stopped writing CID 636 at all for the whole 8-hour verdict — across an overnight charge window. GH#4239's "only retained while a window is configured" is not the whole story — it was refused with slot 1 enabled and its window in force. `read_and_write_cid()` re-reads once after `verify_settle_seconds` before calling a mismatch a failure, because the immediate verify read is taken about half a second after the write. On the SolisCloud select path specifically, **PR #5247 (merged 2026-09-26) makes a discovered inverter's slot-time writes run their handler immediately** so Predbat's read-back sees them in the same cycle — pre-#5247 the select callback only queued the event (GH#4875) and the queue drained at the top of `run()` up to a minute later, so the write verified against the stale value and logged a spurious `didn't complete` once per moved slot time (pass 1, V1 and the slot SoC/current writes keep their every-minute retry). Do not "stop the failed verify poisoning the cache" — `write_cid()` caches the value it *requested* and the post-write `read_cid()` deliberately overwrites it with what the inverter actually reports, so the cache mirrors the inverter rather than Predbat's intent; skipping that update would leave the cache agreeing with the write that just failed, and change detection would then never retry it. The two control paths choose the storage mode differently and it matters: V1 picks from `in_charge_slot`/`in_discharge_slot`, which are clock tests, so CID 636 changes value and gets written the moment a window opens; V2 picks from `slot1_active`, which is only "slot 1 has a window configured", so the mode is written when the slot is *programmed* and nothing at all happens at the window boundary. On the GH#4774 night the last CID 636 write was 22 minutes before the window opened, and the comparison night that worked had one near the window - so `claim_window_mode_assertion()` now asserts the mode once per window on both paths. Slot registers are polled hourly, plus every 5 minutes while `is_inside_active_window()` is true, so drift during the window that matters is caught in the same cycle it happens. GH#4774 also left a useful negative result: that inverter grid-charged from 20% to 91% overnight with bit 1 clear throughout, confirming the per-slot enables alone drive timed charging on V2 — so a "charge window never started" report on V2 firmware is not a CID 636 problem and needs looking at elsewhere. Separately, an inverter reporting `batteryType 'No Battery'` stays in `self.inverter_sn` for polling; `is_battery_inverter()` keeps control writes off it, which was most of the warning volume in GH#4707. That explanation doesn't cover every CID 636 report, though: a later GH#4707 comment showed zero `setting storage mode to` idle-mode log lines, and both V1 mode-decision idle paths always log that line — so a report with neither line cannot be coming from `write_time_windows_if_changed()` at all. The only remaining write path is the HA select handler, `set_storage_mode_value()` at `solis.py:2574`. Check which write path actually ran before assuming this entry's explanation applies. Separately, on inverters using `H M` time format (`GS_fb00`, most cloud subtypes) `adjust_force_export()` used to re-commit a stable export window every cycle — `is_hm_format` alone made `changed_start_end` true, so `press_and_poll_button()` fired on every cycle of an unchanged export window (GH#4709, confirmed live on main at the time) — and PR #4713 fixed the twin idle-cycle press where the times being managed came back as `None` and never compared equal to what the inverter still reported (GH#4712, #2328). **Both merged together in PR #4711 (`bc853a0e`, 2026-09-12, post-v9.0.2)**: `adjust_force_export()` re-commits only on a real schedule change through the commit-once ledger (`last_committed`/`commit_pending`, `commit_needed()`/`record_commit()` in `inverter.py` — the attribute this entry used to name, `last_export_schedule_committed`, is gone), and a commit is only recorded when the writes *and* the button press both succeeded; **the guard only became real when PR #5126 (merged 2026-09-26, unreleased at the time of writing) stopped rebuilding Inverter objects every cycle** — pre-#5126 the rebuild reset the latch to empty each cycle, so the guard was decorative and the every-cycle export commit (the #2328 ~288 TOU-register-write-batches/day signature) stayed live; keep that reading for pre-#5126 logs, where counting `Successfully pressed button` lines settles the actual volume. (The every-cycle `H M` time-entity rewrite remains, deliberately — #1529 write reliability); `press_and_poll_button()` takes a `side` argument so a split-button config no longer presses the unrelated side's button (which is what cleared `timed_charge_current` alongside the intended one); and `adjust_inverter_mode()` now sleeps 30s after a real window change as a GivTCP settle workaround, so expect a 30s pause per real window change on those types. Keep the mechanism for pre-merge logs, where the signature is a button press logged every 5 minutes inside a window. Away from the CID work, `automatic_config()` had an asymmetric arg-binding defect, **fixed in PR #5308 (merged 2026-09-29, v9.3.3) - keep the pre-#5308 signature:** it bound `battery_rate_max` - which fed *both* the charge and the discharge cap - to the per-device `max_charge_power` entity, while the `max_discharge_power` entity it also creates was bound to nothing; `inverter.py`'s `min(inverter_limit_discharge, battery_rate_max_raw)` then capped discharge at the charge limit on an asymmetric inverter, and `inverter_limit_discharge` could not lift it because the raw value was the binding constraint (GH#4940, confirmed against the reporter's debug yaml). The fix: Solis publishes a `battery_rate_max` sensor holding the **larger** of the two per-direction limits and binds `battery_rate_max` to it - auto-discovery still beats a manual apps.yaml value for that binding (`component_base.py:96-110`), so an apps.yaml `battery_rate_max` does not opt out - and binds `inverter_limit_charge`/`inverter_limit_discharge` to the two per-direction limits with `overwrite=False`, so an apps.yaml `inverter_limit_*` (an AC rating, a DNO cap) still wins. Separately, on Solis Cloud/TOU-V2 the per-slot current write is clamped by `write_time_windows()` (`solis.py`, `new_current = min(slot_current, max_discharge_current_amps)`), where the cap is the *minimum* `sysCommand.max` metadata across **all six** slot CIDs from `cached_infos` — deliberately, because a slot 1 value above the others is not accepted. On GH#5068 that bound slot 1 at 60A against a ~93A plan, and the trap is that the number entity is republished from the local `charge_discharge_time_windows` cache, so the write verify reads back the *clamped* value converted at the **live** battery voltage (60A × ~51.6V ≈ 3094W, not 48V) and logs `didn't complete got 3094.0` every cycle — while a "Wrote N successfully now N" line can also be self-referential (cached 60 == clamped 60 skips the CID write entirely). Grep `CID 5967 ... is set to` to see what actually went to the inverter. The planner models `battery_rate_max_discharge` from apps.yaml/entity max and knows nothing of the clamp, so the plan is systematically optimistic while it binds; workaround is `inverter_limit_discharge` at 60A × pack voltage. #4220 is the same 60A clamp from the user's angle. Separately it built every arg list - PV, load and grid included - from the battery-filtered `devices` list (`solis.py:1456`, narrowed by `cd6e7c79`), dropping a PV-only inverter's generation from `pv_today`/`pv_power`; the PV half was fixed in PR #4923, load/grid deliberately left battery-only because those registers can overlap on a shared-CT install. **A string inverter that declares its no-battery state in *neither* place was still enrolled as a battery inverter (GH#5279, fixed in PR #5281, merged 2026-09-28 — keep the pre-#5281 signature):** `_reports_no_battery()` was deliberately narrow (absence of the fields = "unknown", never dropped), and the captured string model (product `0106`) reports `batteryType '0'` — the code batteryList uses for "No Battery" on the alternative firmware, not a name — with an empty `batteryList` and zero battery readings, while `batteryHealthSoh: 0` parses as `0.0` rather than `None`, so `automatic_config()` counted it as a battery and bound control args to it (the #4707 refused-write stream). Signature: `Configuring Predbat for N inverter(s) with batteries` with **no** `Skipping inverter … reports no battery` line and **no** `Including N inverter(s) with no battery in the PV totals` line, plus `Set arg soc_percent = [two entries]`. #5281's detection requires all of: no battery type named (absent, empty or the code `0`), no batteryList entry, and `batteryVoltage` **and** `batteryCapacitySoc` both present and 0 — an unread detail or a named pack reporting zeros still counts as a battery, and a named pack is decided before the readings are looked at; a PV-only inverter now gets `inverterDetail` only (startup TOU read and hourly register reads skipped, publish stops at the detail sensors, event handlers ignore it) and takes its place in the PV totals. The #4923 PV-args split can now apply to such an inverter, which is exactly why it could not before: it never left the battery list. Pre-#5281 workaround remains `solis_inverter_sn` pinned to the battery inverter — which silently removes all PV generation from the plan, so on a current version the fix is the upgrade. One more shape worth remembering: the entity-event queue used to drain at the top of `run()`, before the `first` block created the `ClientSession` and discovered inverters, so an event queued during startup executed against `session=None` with an empty `inverter_sn` - and was popped before execution, making it a silently lost write (fixed in `9fe1f7e0`). Three September fixes and one still-live trap. **PR #5089 (merged 2026-09-14)** closed the SolisCloud API-allowance family (GH#5087/#5091): `B0115` ("Datalogger offline or disconnected") and `R0000` ("Daily API request allowance exhausted") are now classified in `SOLIS_API_CODES` and excluded from retry (`SOLIS_API_CODES_NO_RETRY`) - retrying cannot change the answer and every attempt still spends one of the 200 daily requests - with a per-inverter datalogger-offline cooldown (`datalogger_offline()`) so one offline datalogger backs off alone rather than the fleet; and `automatic_config()` is no longer first-cycle-only: `run()` retries it every cycle until it returns True (`automatic_config_done`), so an API outage at startup no longer leaves `load_today` unset until a Predbat restart (the ValueError-at-fetch symptom this produced has its own row in the symptom table). Keep the mechanism for pre-#5089 logs: B0115 was treated as rate limiting (10s sleep + retry, the `ad26f95a` "Quick rate limit hack") and any R-series code fell through to `Unknown code` with the full retry ladder - an offline datalogger drained the 200/day allowance and R0000 arrived hours later. `automatic_config()` makes no API calls of its own - it reads only `self.inverter_details` - and returns False when details are still missing, which is the retryable case. **PR #5095 (merged 2026-09-15)** fixed the amp↔watt conversion voltage (GH#5090): `get_nominal_voltage()` now prefers `solis_nominal_voltage` from apps.yaml (the only stated physical property), then a live reading above `SOLIS_HV_BATTERY_VOLTAGE` (HV packs keep the live voltage as they have since #4493), then BMS charge voltage classification (16S at/above 55V charge voltage, 15S below - settings rather than measurements, so the result holds across polls), then the 48V fallback; the live voltage had swung 10-14% over a month on every sampled system, dragging `battery_rate_max`, the write tolerance computed from it and the read-back of every rate setpoint along. No workaround existed pre-#5095 - `solis_nominal_voltage` fed only the capacity path (`get_capacity_voltage()`), so telling a user to set it fixed capacity but not the rate paths. Diagnostic tell for pre-#5095 logs: failed-write readbacks at a constant ratio of the target (e.g. exactly 66/70) are consistent with write-time vs read-time voltage difference alone. **Still live (GH#5093):** `SolisAPI.run()` folds per-inverter failures into one `poll_success` bool - discovery failure, any inverter's detail fetch, TOU window reads, the in-window re-read - and `ComponentBase.start()` gates startup on that return: False → `api_started` never set, and the startup burst (discovery, details, `startup_reset_registers()`, the infrequent poll) re-runs on a doubling backoff from 60s to a 128-minute cap, so a multi-inverter fleet is held hostage to the least reachable device and every retry re-spends quota. Per-inverter success booleans already exist at every `poll_success` site, so the fix is contained to `run()`; the interim workaround is listing only working SNs in `solis_inverter_sn`. PR #5185 (merged 2026-09-23) also backs off `startup_reset_registers()` per-inverter with a persistent B0600 pause, and a refused startup register read no longer aborts startup — the try/except now sitting at the `run()` call site covers both triggers, the #5177 B0600 refusal and the SolisCloud timeout (GH#5202, closed as its dup: the traceback matched main line-for-line, the same `SolisAPIError` escaping the same call). Three Modbus-vs-component bindings and a mode-domain split (GH#5127/#5128/#5129, triaged 2026-09-17). On `charge_time_format: "H M"` types (`GS_fb00`, most cloud subtypes) `adjust_charge_window()` rewrites the charge start/end time entities on every 5-minute cycle of a programmed window (the format check in its write condition, inverter.py) - deliberately, the same #1529 write-reliability workaround the export side has - so "Predbat hammers my Solis with writes" is this first, not SoC-hold micromanagement: `adjust_battery_target()` is change-gated (logs `already at target` and writes nothing when equal) and the hold path's `disable_charge_window()` writes nothing unless the enable switch was on. The Modbus-path types (`GS`, `GS_fb00`, `SX4`, Sofar) have **no Predbat component at all** - every control arg is bound only by the hand-edited template; `create_missing_arg()` only ever *dummies* absent capabilities (a dummy `charge_limit` 100 when `has_target_soc` is False), so no INVERTER_DEF default binds a real entity, and a GS_fb00 user who misses the template's `charge_limit` uncomment gets a silent 100.0 target plus `No entity_id for charge_limit to write N` on every target change while the inverter keeps charging to its own default (GH#5129 - the signature is the repeated `No entity_id` line, not a startup error). On the Solis Cloud component the opposite applies: `automatic_config()` binds `charge_limit` with `set_arg_auto(..., overwrite=True)` (solis.py), so an apps.yaml `charge_limit:` pointing at the inverter's own timed-charge-SoC entity is replaced and logged `auto-discovery wins for this setting` - the first thing to grep a reporter's log for when #2328's guidance "didn't help" (GH#5127). Separately, Solis has **two mode-switch control paths whose mode domains do not overlap** (GH#5128): the HA `solax_modbus` integration (`inverter_type: "GS"`) writes a mode *name* from `SOLAX_SOLIS_MODES`/`SOLAX_SOLIS_MODES_NEW` through `alt_charge_discharge_enable()` and only ever computes 33/35 - the 64/96/98 "Feed-in priority" entries are dead in that table - while the SolisCloud API path above does use feed-in priority in its idle paths; so a "why doesn't Predbat use feed-in priority" report is GH#5128's gap on modbus and already-covered on the API component. Trap for any fix there: 33/35 mean different things in the two tables (`SOLAX_SOLIS_MODES` vs `_NEW`, the `solax_modbus_new` switch), so writes must go through the table-derived name, never a bare number; the reusable pattern is `support_feedin_first` (PR #5038, Fox). Check `fox.py`/`gecloud.py` for the same fold-everything-into-one-bool shape - GE Cloud still runs its own auto-config under `if first:` only, with no retry, so its self-heal after an outage is the pre-#5089-era slow one. **GH#5187, fixed in PR #5189 (merged 2026-09-24) - keep the mechanism for pre-#5189 logs.** Slot currents are now capped at the inverter's rated power (from the `inverter_size` sensor, logged `Capping slot currents on ... at the most the inverter can deliver at its rated power`); pre-#5189 the bound was only CID 7226 × count and the per-slot `sysCommand.max` metadata, so a 3.6 kW AC inverter reading 100 A at CID 7226 refused every 100 A slot-current write and left the slot enabled at 0 A, while the planner side never saw the AC rating - pre-#5308 auto-config bound `battery_rate_max` to `max_charge_power` (GH#4940) and set no `inverter_limit_discharge`; since PR #5308 the component binds `inverter_limit_discharge` to `max_discharge_power` and an apps.yaml value still wins (`overwrite=False`), so `inverter_limit_discharge: 3600` remains the plan-side workaround where the modelling matters. Implausible recovery SoC (CID 7229) no longer defeats the #4706 clamp: 0/1/≤over-discharge/>100 now falls back to `over_discharge_soc + 1` with a log line; pre-#5189 a 65521 read passed the guard and made Predbat attempt to write the recovery register itself down every cycle. **Still live:** the write-skip on a never-read slot current CID — `cached_values.get(cid, cap)` defaults to the cap, so a register the poll has never returned reads as "no change" and **no write is attempted at all** (`solis.py`, both charge and discharge paths); if a fixture shows no write of the slot CID, check whether the fixture seeded it in `cached_values` before assuming a bug. Test notes from #5177: `test_solis.py`'s first-cycle tests mock `startup_reset_registers` via instance attribute (`del api.startup_reset_registers` restores the real method); a canned `api.session` set before `run(0, True)` reaches the real request path; and do not patch `asyncio.sleep` to speed up a retried error — `_with_retry` budgets with `time.monotonic()`, so a no-op sleep just burns the real ~30 s budget across ~8 attempts. The hold lever on Solis modbus is window-scoped (GH#5152, code-verified): the only discharge-suppression write is `set_current_from_power()`'s `timed_{charge,discharge}_current` (re-asserted on every rate adjustment, #4415); the standing `battery_discharge_current` registers are never driven anywhere in the repo, and GS makes the timed write the *only* lever (`has_reserve_soc: False`, `has_timed_pause: False`) — so a hold's 0A binds only inside the timed slot, and a "hold/freeze doesn't stop discharge" report on Solis-modbus points at GH#5152 before any plan logic: the execute-side hold gates fire correctly, the lever they pull is what is window-scoped. (The "only binds during the slot" firmware semantics is read off the register names in the solax_modbus integration, not yet live-log confirmed.) Same firmware on the **solax_modbus** path: `inverter_type: GS` with the integration's plain "Solis" plugin on V2/FB00 firmware shows as `Setting Solis Energy Control Switch to 35 Self-Use from 33 Self-Use - No Timed Charge/Discharge` every cycle, the write "succeeds", and the next read is 33 again — a `Control interference: N change(s) in the last 24h, sustained on [...energy_storage_control_switch]` count in the hundreds. Charging still worked (SoC 28→100% overnight) but exports never ran: the discharge slot's times were written but its per-slot enable (43707) stayed 0, which the `solis.py` V2 time-window decode showed as `discharge_enable: 0`. Not a code bug — the user needs the solax_modbus "Solis FB00" plugin (per-slot enable switches, no mode 35) and `inverter_type: GS_fb00` via `templates/ginlong_solis_fb00.yaml`; confirmed fixed on a live system on 26 Sep 2026. Adding Solis Cloud credentials *without* `solis_automatic: True` on such a system leaves inverter 0 bound as GS while `solis.py` still runs its slot/mode logic with no plan — it logs `no active slot`, clears discharge slot 1 and sets 'Self-Use - No Timed Charge/Discharge', fighting the modbus path. **GH#5190 (verified on main): PEP 515 underscores defeat every numeric-parse guard.** Python ≥3.6 `int()`/`float()` accept `_` digit separators, so a junk register string such as `"1234567890_00"` (datalogger off the remote-control platform; polls fail with `B0063 "Lack of iot platform three elements"`) parses as `123456789000` with no exception — `parse_cid_int()`, `cid_value_matches()` and the raw `int(...)`/`float(...)` coercion all accept it, so every `try/except ValueError` fallback silently passes. Consequences seen in the reporter's config: `reserve`/`battery_min_soc` auto-bound to the junk percent (inverter.py raises `set_reserve_min` to it), CID 636 never converging (every-minute write + failed-verify repeat), and no `B0063` handling anywhere in `solis.py`. The same gap is in the core `get_arg()` float coercion (`userinterface.py`), so it generalises past Solis: start there on any "value reads as an absurdly huge number / the fallback never fires" report. Fix shape per the issue: one strict-parse helper (plain decimal only), range checks, publish `None` for unparseable reads, never write CID 636 from an unreadable current mode. Holds and freeze (PR #5265): `reserve` is now bound to the **Battery Reserve SOC** (CID 157, `reserve_soc`) and the row declares `has_reserve_soc` True - `battery_min_soc` stays the over-discharge SOC (CID 158). Solis only treats the Battery Reserve SOC as a discharge floor while CID 636 bit 4 (`SOLIS_BIT_BACKUP_MODE`) is on, so `set_storage_mode_if_needed()` always sets that bit: `set_reserve_enable` only decides whether Predbat *changes* the reserve, and with it off `inverter.py` plans against the live reserve, so the inverter must enforce it too (gating the bit on `set_reserve_enable` left the plan flooring at a register the inverter ignored). Reserve writes are shown in the cache and entity at once by `show_soc_limit_now()` and written from the event queue - non-slot events wait for the next `run()`, far beyond Predbat's 2s write verify. Hold detection has a 1% hysteresis (`SOLIS_HOLD_HYSTERESIS`) so the mode does not flip across the target; a hold outside a slot needs no extra grid-charging logic because outside slots the component already writes a No-Grid-Charging mode. Inside an open charge slot with the battery at or above the slot's SOC (`is_holding_in_charge_slot()` - a freeze charge, or a charge that reached its target) it writes `Self-Use - No Grid Charging` (`ENUM_SELF_USE_HOLD`: 3 on V1 where bit 1 is the slot enable, 1 on V2) ahead of the 0A -> Feed-in rule, which on V1 used to select Feed-in priority - exporting PV - during a freeze charge. | `solis` |
| Solis (`solis.py`) | `SOLIS_CID_STORAGE_MODE = 636` is Modbus 43110, a bit mask. On firmware "4B and above" (what `is_tou_v2_mode()` detects: CID 6798 reads 43605) the timed charge/discharge enable moved to the per-slot registers — `SOLIS_CID_CHARGE_ENABLE_BASE`/`..._DISCHARGE_ENABLE_BASE`, Modbus 43707, six slots not three — and every mode value carrying TOU bit 1 (3/35/43/51/98) was dropped; 35 became 33, 98 became 96. Such an inverter answers a CID 636 write with code 0 and reads back without bit 1, so a log full of CID 636 verification warnings is a refused bit, not a failed write (GH#4707). `set_storage_mode_if_needed()` decides the bit from `is_tou_v2_mode()`: never asked for on V2, always asked for and retried on V1, where bit 1 still *is* the timed charge/discharge enable. Do not try to learn it from the read-back instead — GH#4710 did, and GH#4774 showed why that cannot work: on six inverters the same write of 179 verifies minutes before and minutes after the one that reads back 177, so a single post-write read is not evidence of a firmware property. Latching it was also silent rather than noisy: the verify read refreshes `cached_values`, so after a stripped write the cache held the stripped value, the suppressed computed value matched it, and `set_storage_mode_if_needed()` stopped writing CID 636 at all for the whole 8-hour verdict — across an overnight charge window. GH#4239's "only retained while a window is configured" is not the whole story — it was refused with slot 1 enabled and its window in force. `read_and_write_cid()` re-reads once after `verify_settle_seconds` before calling a mismatch a failure, because the immediate verify read is taken about half a second after the write. On the SolisCloud select path specifically, **PR #5247 (merged 2026-09-26) makes a discovered inverter's slot-time writes run their handler immediately** so Predbat's read-back sees them in the same cycle — pre-#5247 the select callback only queued the event (GH#4875) and the queue drained at the top of `run()` up to a minute later, so the write verified against the stale value and logged a spurious `didn't complete` once per moved slot time (pass 1, V1 and the slot SoC/current writes keep their every-minute retry). Do not "stop the failed verify poisoning the cache" — `write_cid()` caches the value it *requested* and the post-write `read_cid()` deliberately overwrites it with what the inverter actually reports, so the cache mirrors the inverter rather than Predbat's intent; skipping that update would leave the cache agreeing with the write that just failed, and change detection would then never retry it. The two control paths choose the storage mode differently and it matters: V1 picks from `in_charge_slot`/`in_discharge_slot`, which are clock tests, so CID 636 changes value and gets written the moment a window opens; V2 picks from `slot1_active`, which is only "slot 1 has a window configured", so the mode is written when the slot is *programmed* and nothing at all happens at the window boundary. On the GH#4774 night the last CID 636 write was 22 minutes before the window opened, and the comparison night that worked had one near the window - so `claim_window_mode_assertion()` now asserts the mode once per window on both paths. Slot registers are polled hourly, plus every 5 minutes while `is_inside_active_window()` is true, so drift during the window that matters is caught in the same cycle it happens. GH#4774 also left a useful negative result: that inverter grid-charged from 20% to 91% overnight with bit 1 clear throughout, confirming the per-slot enables alone drive timed charging on V2 — so a "charge window never started" report on V2 firmware is not a CID 636 problem and needs looking at elsewhere. Separately, an inverter reporting `batteryType 'No Battery'` stays in `self.inverter_sn` for polling; `is_battery_inverter()` keeps control writes off it, which was most of the warning volume in GH#4707. That explanation doesn't cover every CID 636 report, though: a later GH#4707 comment showed zero `setting storage mode to` idle-mode log lines, and both V1 mode-decision idle paths always log that line — so a report with neither line cannot be coming from `write_time_windows_if_changed()` at all. The only remaining write path is the HA select handler, `set_storage_mode_value()` at `solis.py:2574`. Check which write path actually ran before assuming this entry's explanation applies. Separately, on inverters using `H M` time format (`GS_fb00`, most cloud subtypes) `adjust_force_export()` used to re-commit a stable export window every cycle — `is_hm_format` alone made `changed_start_end` true, so `press_and_poll_button()` fired on every cycle of an unchanged export window (GH#4709, confirmed live on main at the time) — and PR #4713 fixed the twin idle-cycle press where the times being managed came back as `None` and never compared equal to what the inverter still reported (GH#4712, #2328). **Both merged together in PR #4711 (`bc853a0e`, 2026-09-12, post-v9.0.2)**: `adjust_force_export()` re-commits only on a real schedule change through the commit-once ledger (`last_committed`/`commit_pending`, `commit_needed()`/`record_commit()` in `inverter.py` — the attribute this entry used to name, `last_export_schedule_committed`, is gone), and a commit is only recorded when the writes *and* the button press both succeeded; **the guard only became real when PR #5126 (merged 2026-09-26, unreleased at the time of writing) stopped rebuilding Inverter objects every cycle** — pre-#5126 the rebuild reset the latch to empty each cycle, so the guard was decorative and the every-cycle export commit (the #2328 ~288 TOU-register-write-batches/day signature) stayed live; keep that reading for pre-#5126 logs, where counting `Successfully pressed button` lines settles the actual volume. (The every-cycle `H M` time-entity rewrite remains, deliberately — #1529 write reliability); `press_and_poll_button()` takes a `side` argument so a split-button config no longer presses the unrelated side's button (which is what cleared `timed_charge_current` alongside the intended one); and `adjust_inverter_mode()` now sleeps 30s after a real window change as a GivTCP settle workaround, so expect a 30s pause per real window change on those types. Keep the mechanism for pre-merge logs, where the signature is a button press logged every 5 minutes inside a window. Away from the CID work, `automatic_config()` had an asymmetric arg-binding defect, **fixed in PR #5308 (merged 2026-09-29, v9.3.3) - keep the pre-#5308 signature:** it bound `battery_rate_max` - which fed *both* the charge and the discharge cap - to the per-device `max_charge_power` entity, while the `max_discharge_power` entity it also creates was bound to nothing; `inverter.py`'s `min(inverter_limit_discharge, battery_rate_max_raw)` then capped discharge at the charge limit on an asymmetric inverter, and `inverter_limit_discharge` could not lift it because the raw value was the binding constraint (GH#4940, confirmed against the reporter's debug yaml). The fix: Solis publishes a `battery_rate_max` sensor holding the **larger** of the two per-direction limits and binds `battery_rate_max` to it - auto-discovery still beats a manual apps.yaml value for that binding (`component_base.py:96-110`), so an apps.yaml `battery_rate_max` does not opt out - and binds `inverter_limit_charge`/`inverter_limit_discharge` to the two per-direction limits with `overwrite=False`, so an apps.yaml `inverter_limit_*` (an AC rating, a DNO cap) still wins. Separately, on Solis Cloud/TOU-V2 the per-slot current write is clamped by `write_time_windows()` (`solis.py`, `new_current = min(slot_data['discharge_current'], cap)`). **The cap's source was rewritten in PR #5314 (merged 2026-09-30, v9.3.3) - keep the pre-#5314 signature:** the cap used to be the *minimum* `sysCommand.max` metadata across all six slot CIDs from `cached_infos` — read here at the time as deliberate, but the metadata is actually a UI definition for a *list of model codes* (its `productModel`), and the cloud hands back definitions for other models: an S5-EH1P5K-L (model 3104, 5kW, registers reading 100A) was given 60A for its slot currents from the models-3101/3102 definition and 62.5A for slot 1 charge from models 3111/3140/3145's, holding exports to ~3.6kW (GH#5068). `slot_current_limits()` now ignores the metadata: the cap starts from the battery-side limit register itself (CID 7224 charge / 7226 discharge, read as-is - not multiplied by pack count - and `SOLIS_SLOT_CURRENT_DEFAULT_AMPS` when never read), and once `ensure_slot_current_probed()` has measured what the slots really accept a **measured** ceiling replaces the rated estimate (PR #5309's probe, runnable as `python3 apps/predbat/solis.py --probe-ceiling` without `--write`; it starts from the register and the sweep ends with a half-amp step, so a 62.5A ceiling is found exactly rather than rounded down, with candidates written as the inverter shows them). The verify-echo traps survive the rewrite: the number entity is republished from the local `charge_discharge_time_windows` cache, so the write verify reads back the *clamped* value converted at the **live** battery voltage (60A × ~51.6V ≈ 3094W, not 48V) and logs `didn't complete got 3094.0` every cycle — while a "Wrote N successfully now N" line can also be self-referential (cached == clamped skips the CID write entirely). Grep `CID 5967 ... is set to` to see what actually went to the inverter. The planner models `battery_rate_max_discharge` from apps.yaml/entity max and knows nothing of the clamp, so the plan is systematically optimistic while it binds; on a current build set `inverter_limit_discharge` to the measured cap × pack voltage, pre-#5314 it was 60A × pack voltage. #4220 is the same 60A clamp from the user's angle. Separately it built every arg list - PV, load and grid included - from the battery-filtered `devices` list (`solis.py:1456`, narrowed by `cd6e7c79`), dropping a PV-only inverter's generation from `pv_today`/`pv_power`; the PV half was fixed in PR #4923, load/grid deliberately left battery-only because those registers can overlap on a shared-CT install. **A string inverter that declares its no-battery state in *neither* place was still enrolled as a battery inverter (GH#5279, fixed in PR #5281, merged 2026-09-28 — keep the pre-#5281 signature):** `_reports_no_battery()` was deliberately narrow (absence of the fields = "unknown", never dropped), and the captured string model (product `0106`) reports `batteryType '0'` — the code batteryList uses for "No Battery" on the alternative firmware, not a name — with an empty `batteryList` and zero battery readings, while `batteryHealthSoh: 0` parses as `0.0` rather than `None`, so `automatic_config()` counted it as a battery and bound control args to it (the #4707 refused-write stream). Signature: `Configuring Predbat for N inverter(s) with batteries` with **no** `Skipping inverter … reports no battery` line and **no** `Including N inverter(s) with no battery in the PV totals` line, plus `Set arg soc_percent = [two entries]`. #5281's detection requires all of: no battery type named (absent, empty or the code `0`), no batteryList entry, and `batteryVoltage` **and** `batteryCapacitySoc` both present and 0 — an unread detail or a named pack reporting zeros still counts as a battery, and a named pack is decided before the readings are looked at; a PV-only inverter now gets `inverterDetail` only (startup TOU read and hourly register reads skipped, publish stops at the detail sensors, event handlers ignore it) and takes its place in the PV totals. The #4923 PV-args split can now apply to such an inverter, which is exactly why it could not before: it never left the battery list. Pre-#5281 workaround remains `solis_inverter_sn` pinned to the battery inverter — which silently removes all PV generation from the plan, so on a current version the fix is the upgrade. One more shape worth remembering: the entity-event queue used to drain at the top of `run()`, before the `first` block created the `ClientSession` and discovered inverters, so an event queued during startup executed against `session=None` with an empty `inverter_sn` - and was popped before execution, making it a silently lost write (fixed in `9fe1f7e0`). Three September fixes and one still-live trap. **PR #5089 (merged 2026-09-14)** closed the SolisCloud API-allowance family (GH#5087/#5091): `B0115` ("Datalogger offline or disconnected") and `R0000` ("Daily API request allowance exhausted") are now classified in `SOLIS_API_CODES` and excluded from retry (`SOLIS_API_CODES_NO_RETRY`) - retrying cannot change the answer and every attempt still spends one of the 200 daily requests - with a per-inverter datalogger-offline cooldown (`datalogger_offline()`) so one offline datalogger backs off alone rather than the fleet; and `automatic_config()` is no longer first-cycle-only: `run()` retries it every cycle until it returns True (`automatic_config_done`), so an API outage at startup no longer leaves `load_today` unset until a Predbat restart (the ValueError-at-fetch symptom this produced has its own row in the symptom table). Keep the mechanism for pre-#5089 logs: B0115 was treated as rate limiting (10s sleep + retry, the `ad26f95a` "Quick rate limit hack") and any R-series code fell through to `Unknown code` with the full retry ladder - an offline datalogger drained the 200/day allowance and R0000 arrived hours later. `automatic_config()` makes no API calls of its own - it reads only `self.inverter_details` - and returns False when details are still missing, which is the retryable case. **PR #5095 (merged 2026-09-15)** fixed the amp↔watt conversion voltage (GH#5090): `get_nominal_voltage()` now prefers `solis_nominal_voltage` from apps.yaml (the only stated physical property), then a live reading above `SOLIS_HV_BATTERY_VOLTAGE` (HV packs keep the live voltage as they have since #4493), then BMS charge voltage classification (16S at/above 55V charge voltage, 15S below - settings rather than measurements, so the result holds across polls), then the 48V fallback; the live voltage had swung 10-14% over a month on every sampled system, dragging `battery_rate_max`, the write tolerance computed from it and the read-back of every rate setpoint along. No workaround existed pre-#5095 - `solis_nominal_voltage` fed only the capacity path (`get_capacity_voltage()`), so telling a user to set it fixed capacity but not the rate paths. Diagnostic tell for pre-#5095 logs: failed-write readbacks at a constant ratio of the target (e.g. exactly 66/70) are consistent with write-time vs read-time voltage difference alone. **Still live (GH#5093):** `SolisAPI.run()` folds per-inverter failures into one `poll_success` bool - discovery failure, any inverter's detail fetch, TOU window reads, the in-window re-read - and `ComponentBase.start()` gates startup on that return: False → `api_started` never set, and the startup burst (discovery, details, `startup_reset_registers()`, the infrequent poll) re-runs on a doubling backoff from 60s to a 128-minute cap, so a multi-inverter fleet is held hostage to the least reachable device and every retry re-spends quota. Per-inverter success booleans already exist at every `poll_success` site, so the fix is contained to `run()`; the interim workaround is listing only working SNs in `solis_inverter_sn`. PR #5185 (merged 2026-09-23) also backs off `startup_reset_registers()` per-inverter with a persistent B0600 pause, and a refused startup register read no longer aborts startup — the try/except now sitting at the `run()` call site covers both triggers, the #5177 B0600 refusal and the SolisCloud timeout (GH#5202, closed as its dup: the traceback matched main line-for-line, the same `SolisAPIError` escaping the same call). Three Modbus-vs-component bindings and a mode-domain split (GH#5127/#5128/#5129, triaged 2026-09-17). On `charge_time_format: "H M"` types (`GS_fb00`, most cloud subtypes) `adjust_charge_window()` rewrites the charge start/end time entities on every 5-minute cycle of a programmed window (the format check in its write condition, inverter.py) - deliberately, the same #1529 write-reliability workaround the export side has - so "Predbat hammers my Solis with writes" is this first, not SoC-hold micromanagement: `adjust_battery_target()` is change-gated (logs `already at target` and writes nothing when equal) and the hold path's `disable_charge_window()` writes nothing unless the enable switch was on. The Modbus-path types (`GS`, `GS_fb00`, `SX4`, Sofar) have **no Predbat component at all** - every control arg is bound only by the hand-edited template; `create_missing_arg()` only ever *dummies* absent capabilities (a dummy `charge_limit` 100 when `has_target_soc` is False), so no INVERTER_DEF default binds a real entity, and a GS_fb00 user who misses the template's `charge_limit` uncomment gets a silent 100.0 target plus `No entity_id for charge_limit to write N` on every target change while the inverter keeps charging to its own default (GH#5129 - the signature is the repeated `No entity_id` line, not a startup error). On the Solis Cloud component the opposite applies: `automatic_config()` binds `charge_limit` with `set_arg_auto(..., overwrite=True)` (solis.py), so an apps.yaml `charge_limit:` pointing at the inverter's own timed-charge-SoC entity is replaced and logged `auto-discovery wins for this setting` - the first thing to grep a reporter's log for when #2328's guidance "didn't help" (GH#5127). Separately, Solis has **two mode-switch control paths whose mode domains do not overlap** (GH#5128): the HA `solax_modbus` integration (`inverter_type: "GS"`) writes a mode *name* from `SOLAX_SOLIS_MODES`/`SOLAX_SOLIS_MODES_NEW` through `alt_charge_discharge_enable()` and only ever computes 33/35 - the 64/96/98 "Feed-in priority" entries are dead in that table - while the SolisCloud API path above does use feed-in priority in its idle paths; so a "why doesn't Predbat use feed-in priority" report is GH#5128's gap on modbus and already-covered on the API component. Trap for any fix there: 33/35 mean different things in the two tables (`SOLAX_SOLIS_MODES` vs `_NEW`, the `solax_modbus_new` switch), so writes must go through the table-derived name, never a bare number; the reusable pattern is `support_feedin_first` (PR #5038, Fox). Check `fox.py`/`gecloud.py` for the same fold-everything-into-one-bool shape - GE Cloud still runs its own auto-config under `if first:` only, with no retry, so its self-heal after an outage is the pre-#5089-era slow one. **GH#5187, fixed in PR #5189 (merged 2026-09-24) - keep the mechanism for pre-#5189 logs.** Slot currents are now capped at the inverter's rated power (from the `inverter_size` sensor, logged `Capping slot currents on ... at the most the inverter can deliver at its rated power`); pre-#5189 the bound was only CID 7226 × count and the per-slot `sysCommand.max` metadata, so a 3.6 kW AC inverter reading 100 A at CID 7226 refused every 100 A slot-current write and left the slot enabled at 0 A, while the planner side never saw the AC rating - pre-#5308 auto-config bound `battery_rate_max` to `max_charge_power` (GH#4940) and set no `inverter_limit_discharge`; since PR #5308 the component binds `inverter_limit_discharge` to `max_discharge_power` and an apps.yaml value still wins (`overwrite=False`), so `inverter_limit_discharge: 3600` remains the plan-side workaround where the modelling matters. Implausible recovery SoC (CID 7229) no longer defeats the #4706 clamp: 0/1/≤over-discharge/>100 now falls back to `over_discharge_soc + 1` with a log line; pre-#5189 a 65521 read passed the guard and made Predbat attempt to write the recovery register itself down every cycle. **Still live:** the write-skip on a never-read slot current CID — `cached_values.get(cid, cap)` defaults to the cap, so a register the poll has never returned reads as "no change" and **no write is attempted at all** (`solis.py`, both charge and discharge paths); if a fixture shows no write of the slot CID, check whether the fixture seeded it in `cached_values` before assuming a bug. Test notes from #5177: `test_solis.py`'s first-cycle tests mock `startup_reset_registers` via instance attribute (`del api.startup_reset_registers` restores the real method); a canned `api.session` set before `run(0, True)` reaches the real request path; and do not patch `asyncio.sleep` to speed up a retried error — `_with_retry` budgets with `time.monotonic()`, so a no-op sleep just burns the real ~30 s budget across ~8 attempts. The hold lever on Solis modbus is window-scoped (GH#5152, code-verified): the only discharge-suppression write is `set_current_from_power()`'s `timed_{charge,discharge}_current` (re-asserted on every rate adjustment, #4415); the standing `battery_discharge_current` registers are never driven anywhere in the repo, and GS makes the timed write the *only* lever (`has_reserve_soc: False`, `has_timed_pause: False`) — so a hold's 0A binds only inside the timed slot, and a "hold/freeze doesn't stop discharge" report on Solis-modbus points at GH#5152 before any plan logic: the execute-side hold gates fire correctly, the lever they pull is what is window-scoped. (The "only binds during the slot" firmware semantics is read off the register names in the solax_modbus integration, not yet live-log confirmed.) Same firmware on the **solax_modbus** path: `inverter_type: GS` with the integration's plain "Solis" plugin on V2/FB00 firmware shows as `Setting Solis Energy Control Switch to 35 Self-Use from 33 Self-Use - No Timed Charge/Discharge` every cycle, the write "succeeds", and the next read is 33 again — a `Control interference: N change(s) in the last 24h, sustained on [...energy_storage_control_switch]` count in the hundreds. Charging still worked (SoC 28→100% overnight) but exports never ran: the discharge slot's times were written but its per-slot enable (43707) stayed 0, which the `solis.py` V2 time-window decode showed as `discharge_enable: 0`. Not a code bug — the user needs the solax_modbus "Solis FB00" plugin (per-slot enable switches, no mode 35) and `inverter_type: GS_fb00` via `templates/ginlong_solis_fb00.yaml`; confirmed fixed on a live system on 26 Sep 2026. Adding Solis Cloud credentials *without* `solis_automatic: True` on such a system leaves inverter 0 bound as GS while `solis.py` still runs its slot/mode logic with no plan — it logs `no active slot`, clears discharge slot 1 and sets 'Self-Use - No Timed Charge/Discharge', fighting the modbus path. **GH#5190 (verified on main): PEP 515 underscores defeat every numeric-parse guard.** Python ≥3.6 `int()`/`float()` accept `_` digit separators, so a junk register string such as `"1234567890_00"` (datalogger off the remote-control platform; polls fail with `B0063 "Lack of iot platform three elements"`) parses as `123456789000` with no exception — `parse_cid_int()`, `cid_value_matches()` and the raw `int(...)`/`float(...)` coercion all accept it, so every `try/except ValueError` fallback silently passes. Consequences seen in the reporter's config: `reserve`/`battery_min_soc` auto-bound to the junk percent (inverter.py raises `set_reserve_min` to it), CID 636 never converging (every-minute write + failed-verify repeat), and no `B0063` handling anywhere in `solis.py`. The same gap is in the core `get_arg()` float coercion (`userinterface.py`), so it generalises past Solis: start there on any "value reads as an absurdly huge number / the fallback never fires" report. Fix shape per the issue: one strict-parse helper (plain decimal only), range checks, publish `None` for unparseable reads, never write CID 636 from an unreadable current mode. Holds and freeze (PR #5265): `reserve` is now bound to the **Battery Reserve SOC** (CID 157, `reserve_soc`) and the row declares `has_reserve_soc` True - `battery_min_soc` stays the over-discharge SOC (CID 158). Solis only treats the Battery Reserve SOC as a discharge floor while CID 636 bit 4 (`SOLIS_BIT_BACKUP_MODE`) is on, so `set_storage_mode_if_needed()` always sets that bit: `set_reserve_enable` only decides whether Predbat *changes* the reserve, and with it off `inverter.py` plans against the live reserve, so the inverter must enforce it too (gating the bit on `set_reserve_enable` left the plan flooring at a register the inverter ignored). Reserve writes are shown in the cache and entity at once by `show_soc_limit_now()` and written from the event queue - non-slot events wait for the next `run()`, far beyond Predbat's 2s write verify. Hold detection has a 1% hysteresis (`SOLIS_HOLD_HYSTERESIS`) so the mode does not flip across the target; a hold outside a slot needs no extra grid-charging logic because outside slots the component already writes a No-Grid-Charging mode. Inside an open charge slot with the battery at or above the slot's SOC (`is_holding_in_charge_slot()` - a freeze charge, or a charge that reached its target) it writes `Self-Use - No Grid Charging` (`ENUM_SELF_USE_HOLD`: 3 on V1 where bit 1 is the slot enable, 1 on V2) ahead of the 0A -> Feed-in rule, which on V1 used to select Feed-in priority - exporting PV - during a freeze charge. | `solis` |
Comment thread tools/debug-journal.md
@@ -128,15 +128,15 @@ Grep for the named symbol rather than trusting a line number.
| Standalone / Docker (non-HA) | GH#4601: a callback returning `None` instead of `True` broke the Octopus saving-session fallback in standalone mode. Anything that works under HA but not standalone is worth checking along the `ha.py` websocket and `userinterface.py` callback paths. | `trigger_callback_success_signal` |
| Holiday mode (`fetch.py`) | GH#4732: under `days_previous_auto` (the default) holiday mode is handled entirely inside `compute_load_forecast_history()`, not by the `days_previous = [1]` branch, which is only reachable with `days_previous_auto: False`. Days whose holiday state does not match the *forecast day's* are now excluded outright rather than halved, and `get_holiday_minutes()` must span `num_days + 1` to cover the `minutes_now` overhang. A slot with no matching history falls back to `holiday_load_scaling` (default 0.7), which is what makes holiday mode act on day one - a normalised weighted mean cannot otherwise express "all of my data is wrong". | `holiday_mode` |
| Predheat (`predheat.py`) | GH#4670: with `predheat_enable` set, Predheat still did not activate after startup because of lazy flag initialisation. There is no registered Predheat test module, so there is nothing to run here — investigate by reading. GH#4848 mapped the surface for the recurring "make Predheat drive my heat pump" requests: Predheat has **no control path at all** - its only outbound service call is a read-only `weather/get_forecasts`, everything else is `set_state` publishing prediction sensors. `smart_thermostat` looks like control but only pre-empts an already-scheduled setpoint rise *inside the simulation*. And Predheat is instantiated as its own object with its own timer loop rather than as a `PredBat` mixin, so `plan.py` never sees a heat decision variable. `load_forecast: - predheat.heat_energy$external` does surface heat consumption as the plan's Xload column, but as a *fixed* load the battery plans around, never a shifted one. **GH#5153** (code + repro verified): `get_weather_data()` passes the forecast's `temperature` raw into `minute_data()` with no unit conversion — a °F-native weather entity feeds ~55-66 into a °C model, the loss term `heat_loss_watts * (internal - external)` then *heats* the house toward the modelled outside temperature and the internal forecast runs away and levels just below it while energy/cost stay 0 because the thermostat correctly never fires. One-glance check: the weather entity's `temperature_unit` attribute. Same issue, second bug: `minute_data_age` includes the age of `heating_energy`, which is optional — unconfigured, its age is 0, `minute_data_age` pins to 0 and `get_historical()` returns its 20.0 default for the target sensor, so the live target is ignored whenever heating_energy is unset (the template includes it; ASHP users without a heat meter remove it). Traps: `run_simulation(save="best")` publishes HA entities by default, so a parameter sweep overwrites the published sensors unless `save` is set otherwise; predheat reads its own clock (`datetime.now(local_tz)`, not the fixture's pinned `now_utc`), so history mocks must be built against the real wall clock or `minute_data_age` comes out negative; and the registered `test_predheat` stubs `update_pred`, so no simulation physics is under any test. | none |
| Load ML CPU spikes (`load_ml_component.py`, `load_predictor.py`) | The 2-hourly retrain isn't one pass. `_do_training()` hardcodes `ml_curriculum_step_days = 1` / `ml_curriculum_max_passes = 4` (`load_ml_component.py:94-95` — plain instance attributes, no `config.py` entry, not user-configurable), and `train_curriculum()` caps to the largest N windows (`load_predictor.py:1560-1564`); on ~80 days of history that's roughly five near-full-scale passes back to back, confirmed against a reporter's log as a 16-minute CPU spike every 2 hours (GH#3896). Separately, `threads` only ever reaches the C++ prediction kernel (`plan.py:1497`, `resolve_batch_threads`) — never wired to ML — and there is no `OMP_NUM_THREADS`/`threadpoolctl` anywhere in the repo (confirmed absent, 2026-08-27), so NumPy's BLAS backend is free to fan out across every core on its own during that training window. Whether that actually saturates a host depends on which BLAS backend is linked — worth confirming the platform before assuming a code fix is the right lever. There is a mapped test now: `ml_training_perf` (PR #4912), which asserts the training invariant rather than peak RSS — a peak-memory assertion was too machine-dependent to hold. Frequency is a separate axis (GH#5072): `RETRAIN_INTERVAL_SECONDS` (2h, `load_ml_component.py`) is the only knob of the cadence and is hardcoded, and any "make the retrain interval configurable" implementation hits a silent ceiling at the equally hardcoded `ml_max_model_age_hours = 48` — `LoadPredictor.is_valid()` returns invalid with reason `"stale"` past it and ML predictions then fall back to empty, so exposing the interval without also exposing the staleness cap is a footgun above 48h. Model-status signals (GH#5075, code-verified on main; PR #5112 open draft carries the fix): `is_valid()` can report **"active" for a model that has never been trained** — `_initialize_weights()` sets `model_initialized = True` before the first epoch, and both `validation_mae` and the age check skip `None`, so a training abandoned at epoch 0 publishes `model_status = "active"` / `model_valid = True`; conversely `training_timestamp is None` is *not* a sound "never trained" test, because a legacy saved model that predates the field is trained and must stay valid (`test_load_ml.py` asserts exactly that). And `train()` stamps `training_timestamp`/`validation_mae` at the end of *every* curriculum pass, so a curriculum that aborts part-way leaves the last intermediate window's **fresh** stamps standing — suppressing exactly the `ml_max_model_age_hours` staleness retrain that would otherwise rebuild a complete model. #5112 adds a `model_trained` flag plus stamp snapshot/restore around `train_curriculum()`; until it merges, treat "status active + nonsense forecast on a just-started install, or whose initial training keeps failing" as this, not as a data problem. The normalisation-statistics pairing on an abort (restored weights, stats refit for the abandoned run) is *suspected, not measured*. **"Reverting fixed the CPU" is usually an observation-window artefact (GH#5213, verified from a Pi 5 reporter's log):** load-ML fine-tune ran on *both* versions with identical structure (5 passes, 30 epochs, ~14s/epoch — 20–41 min every 2h ≈ 19% duty of one core), while both of the reporter's short v9.0.3 windows contained a training their next upgrade killed partway, so the panel always read low when they checked after rolling back. On any "CPU regressed after upgrade to X" report with Load ML on, first reconstruct from the log when fine-tunes started/finished under each version and how long the user actually stayed on each — the 2h retrain boundary means a version window shorter than that proves nothing. | `ml_training_perf` |
| Load ML (`load_ml_component.py`, `load_predictor.py`) | The 2-hourly retrain isn't one pass. `_do_training()` hardcodes `ml_curriculum_step_days = 1` / `ml_curriculum_max_passes = 4` (`load_ml_component.py:94-95` — plain instance attributes, no `config.py` entry, not user-configurable), and `train_curriculum()` caps to the largest N windows (`load_predictor.py:1560-1564`); on ~80 days of history that's roughly five near-full-scale passes back to back, confirmed against a reporter's log as a 16-minute CPU spike every 2 hours (GH#3896). Separately, `threads` only ever reaches the C++ prediction kernel (`plan.py:1497`, `resolve_batch_threads`) — never wired to ML — and there is no `OMP_NUM_THREADS`/`threadpoolctl` anywhere in the repo (confirmed absent, 2026-08-27), so NumPy's BLAS backend is free to fan out across every core on its own during that training window. Whether that actually saturates a host depends on which BLAS backend is linked — worth confirming the platform before assuming a code fix is the right lever. There is a mapped test now: `ml_training_perf` (PR #4912), which asserts the training invariant rather than peak RSS — a peak-memory assertion was too machine-dependent to hold. Frequency is a separate axis (GH#5072): `RETRAIN_INTERVAL_SECONDS` (2h, `load_ml_component.py`) is the only knob of the cadence and is hardcoded, and any "make the retrain interval configurable" implementation hits a silent ceiling at the equally hardcoded `ml_max_model_age_hours = 48` — `LoadPredictor.is_valid()` returns invalid with reason `"stale"` past it and ML predictions then fall back to empty, so exposing the interval without also exposing the staleness cap is a footgun above 48h. Model-status signals (GH#5075, code-verified on main; PR #5112 open draft carries the fix): `is_valid()` can report **"active" for a model that has never been trained** — `_initialize_weights()` sets `model_initialized = True` before the first epoch, and both `validation_mae` and the age check skip `None`, so a training abandoned at epoch 0 publishes `model_status = "active"` / `model_valid = True`; conversely `training_timestamp is None` is *not* a sound "never trained" test, because a legacy saved model that predates the field is trained and must stay valid (`test_load_ml.py` asserts exactly that). And `train()` stamps `training_timestamp`/`validation_mae` at the end of *every* curriculum pass, so a curriculum that aborts part-way leaves the last intermediate window's **fresh** stamps standing — suppressing exactly the `ml_max_model_age_hours` staleness retrain that would otherwise rebuild a complete model. #5112 adds a `model_trained` flag plus stamp snapshot/restore around `train_curriculum()`; until it merges, treat "status active + nonsense forecast on a just-started install, or whose initial training keeps failing" as this, not as a data problem. The normalisation-statistics pairing on an abort (restored weights, stats refit for the abandoned run) is *suspected, not measured*. **"Reverting fixed the CPU" is usually an observation-window artefact (GH#5213, verified from a Pi 5 reporter's log):** load-ML fine-tune ran on *both* versions with identical structure (5 passes, 30 epochs, ~14s/epoch — 20–41 min every 2h ≈ 19% duty of one core), while both of the reporter's short v9.0.3 windows contained a training their next upgrade killed partway, so the panel always read low when they checked after rolling back. On any "CPU regressed after upgrade to X" report with Load ML on, first reconstruct from the log when fine-tunes started/finished under each version and how long the user actually stayed on each — the 2h retrain boundary means a version window shorter than that proves nothing. **Load ML has a hardcoded 48h horizon and the plan's load is zero beyond it (GH#5336, enhancement, read on main 57ec7bf1 2026-10-01):** `PREDICT_HORIZON = 48 * (60 // CHUNK_MINUTES)` (`load_predictor.py:35`) is the only horizon definition, and it feeds only the `predict()` rollout loop, blend schedule, logs and the `save()` metadata — training does not depend on it, and `load()` validates only model version + architecture, so models survive a horizon change. With `load_ml_source` set, `fetch.py` sets `load_forecast_only = True` and the weighted-bucket historical forecast is skipped wholesale, so the plan's future load is `self.load_forecast` alone (via `step_data_history(..., load_forecast=...)` from `fetch_sensor_data()`'s call path) — past minute 2880 there is no data and the plan-HTML load column is exactly 0; the log tell is `Starting autoregressive prediction loop for 576 steps (48.0 hours)`. The "over 48h" in `Generated N predictions (total X kWh over 48h)` (`load_ml_component.py:695`) is hardcoded in the *log string* — the step count is the truth — and `test_load_ml.py` pins 576 steps. The ask (configurable 48/72/96) is architecturally straightforward because little else depends on the constant, and there is no fallback beyond the ML horizon by design of `load_forecast_only`. | `ml_training_perf` |
Comment thread tools/debug-journal.md
@@ -180,7 +182,7 @@ Grep for the named symbol rather than trusting a line number.
| Works in HA, broken in Docker/standalone | `ha.py` websocket and `userinterface.py` callback return values. |
| Every charge/export window executes ~N minutes late | The inverter's own clock. Compensation is manual-only via the four `inverter_clock_skew_*` settings. The warn band has moved twice: originally only ≥30 minutes was flagged (restart threshold, `INVERTER_CLOCK_SKEW_RESTART_MINUTES`), so a consistent 25-minute drift sailed under it while still logging a `difference -25.47 minutes` line every cycle (GH#4927); since PR #4991 (`87d4703a`, v9.0.2) `check_clock_skew()` also warns on a moderate 5-29 minute skew, repeated at most hourly per inverter (`INVERTER_CLOCK_SKEW_WARN_MINUTES` = 5, `INVERTER_CLOCK_SKEW_WARN_REPEAT_MINUTES` = 60), so a silent moderate skew is no longer the explanation on a current version — but "current" means v9.0.2+: a reporter on v9.0.1 or earlier is back to "only ≥30 minutes was flagged". Measured from one reporter's log: an export window written at 23:05 drew house-load import until 23:35, and a charge window enabled at 00:30 did not start until ~01:00 (GH#4927). |
| SoC reads far too high, or never falls below ~90% | The sensor mapped to `soc_kw`. A cumulative "battery charge today" energy sensor makes SoC ratchet up all day and reset at midnight. Note `battery_soc:` is **not** a config key - only `soc_percent` and `soc_kw` are read, and unknown apps.yaml keys are silently ignored, so a correct-looking `battery_soc:` block mapping the right sensor does nothing at all (GH#4884). |
| An AppDaemon-published sensor disappears after an HA/Predbat restart and never comes back, while planning and values still work | Check whether the sensor's only publish site sits inside the branch that *reads* the recovered data. Instance: `battery_size_tracking()` (`inverter.py`) publishes `sensor.<prefix>_soc_max_calculated` only behind `if today_key not in existing_history:` — and since PR #4259 `existing_history` is recovered from the recorder via `load_previous_value_from_ha()`, so a mid-day restart (HA or AppDaemon — the entity is AppDaemon-published and vanishes on any AppDaemon restart) skips the publish every cycle until midnight (GH#5222). The trimmed mean *is* recovered separately, so `battery_scaling_auto` and planning are unaffected — which is why everything else looks fine. The same failure triggers on a Predbat/AppDaemon restart, not just an HA restart — wider than the issue title implies; other `load_previous_value_from_ha()` callers (savings totals, ohme energy_today) re-publish unconditionally each cycle, so this gated-publish shape is currently unique to `battery_size_tracking`. Test trap: the existing restart-regression test pops the entity and asserts only mean/scaling recovery — never that the entity is re-published, so a green suite says nothing about the publish half. Fix direction (maintainer's call): re-publish from the recovered entry, preserving a recovered `None` rather than recomputing. |
| An AppDaemon-published sensor disappears after an HA/Predbat restart and never comes back, while planning and values still work | Check whether the sensor's only publish site sits inside the branch that *reads* the recovered data. Instance: `battery_size_tracking()` (`inverter.py`) publishes `sensor.<prefix>_soc_max_calculated` only behind `if today_key not in existing_history:` — and since PR #4259 `existing_history` is recovered from the recorder via `load_previous_value_from_ha()`, so a mid-day restart (HA or AppDaemon — the entity is AppDaemon-published and vanishes on any AppDaemon restart) skips the publish every cycle until midnight (GH#5222). The trimmed mean *is* recovered separately, so `battery_scaling_auto` and planning are unaffected — which is why everything else looks fine. The same failure triggers on a Predbat/AppDaemon restart, not just an HA restart — wider than the issue title implies; other `load_previous_value_from_ha()` callers (savings totals, ohme energy_today) re-publish unconditionally each cycle, so this gated-publish shape was unique to `battery_size_tracking`. **Fixed in PR #5283 (merged 2026-09-30, v9.3.3)** — `battery_size_tracking()` now re-publishes the sensor from the recorder-recovered entry when a restart removed the live one, preserving a recovered `None` (today's failed measurement) rather than recomputing, and never re-publishes again once the sensor exists; the restart-regression tests now cover both the republish and the no-repeat half. |
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants