Repository navigation
Fix bootloop caused by orphaned pool workers and solis debug logs - #3718
Conversation
…ception - Remove accidental debug log lines from solis.py read_cid, read_batch, and write_cid methods that were left in from the solis_fix PR (#3716) - Fix misleading description in startup_reset_registers third API call - Add cleanup_pool() helper method to PredBat to properly terminate multiprocessing pool workers - Call cleanup_pool() in reset() to clean up any orphaned pool from a previous failed run on restart - Call cleanup_pool() in update_time_loop and run_time_loop finally blocks to ensure pool workers are terminated when an exception propagates out of the prediction loop (prevents orphaned worker processes that could contribute to the bootloop) Agent-Logs-Url: https://github.com/springfall2008/batpred/sessions/67d7d029-d664-4901-979c-e2b278b7a45b Co-authored-by: springfall2008 <48591903+springfall2008@users.noreply.github.com>
There was a problem hiding this comment.
Pull request overview
Fixes an AppDaemon bootloop/resource exhaustion scenario by ensuring multiprocessing.Pool workers are force-cleaned up when the prediction loop exits unexpectedly, and removes accidentally committed Solis API debug logging (plus a small log-description correction).
Changes:
- Removed verbose Solis debug logs that printed full request/response payloads on every CID read/write.
- Added a
cleanup_pool()helper to terminate/join an activemultiprocessing.Pooland invoked it during reset and after prediction-loop runs. - Corrected
startup_reset_registers’sfield_descriptionto reflect restoring the prior mode.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.
| File | Description |
|---|---|
| apps/predbat/solis.py | Removes production debug logging of Solis API payloads/responses and fixes a misleading field_description string. |
| apps/predbat/predbat.py | Adds centralized pool cleanup via cleanup_pool() and calls it from reset() and prediction-loop finally blocks to prevent orphaned workers. |
| def cleanup_pool(self): | ||
| """ | ||
| Terminate and clean up the multiprocessing pool if it is active. | ||
|
|
||
| Ensures worker processes are properly terminated to prevent orphaned |
There was a problem hiding this comment.
New pool-cleanup behavior is critical to preventing the bootloop, but there doesn’t appear to be any unit test coverage exercising it (e.g., simulating an exception during prediction with a mocked Pool and asserting terminate/join are invoked and self.pool is cleared). Adding a lightweight test that patches the Pool object would help prevent regressions without needing to spawn real worker processes.
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
After the v8.35.2 solis_fix PR, a bootloop was observed where AppDaemon restarts predbat after each prediction run, leaving multiprocessing pool workers orphaned. Additionally, debug log statements were accidentally left in production code.
Changes
solis.py— Remove debug logs# Debug logstatements accidentally committed inread_cid,read_batch, andwrite_cidthat logged full API payloads and responses on every callfield_descriptioninstartup_reset_registers: the thirdread_and_write_cidcall restores the original mode but was labelled identically to the first call that sets the new modepredbat.py— Robust pool cleanupAdded
cleanup_pool()helper that safely terminates and joins any activemultiprocessing.Pool:Called in three places to prevent orphaned worker processes:
reset()— terminates any pool left over from a prior crashed run when the app restartsupdate_time_loop/run_time_loopfinallyblocks — terminates the pool when an exception propagates out ofupdate_predbeforecalculate_plan's normal cleanup (pool.close()+pool.join()) is reachedPreviously, an exception inside the prediction loop would leave
self.poolpointing to live worker processes. On the next AppDaemon restartreset()would silently overwriteself.pool = None, orphaning those workers. On constrained hardware this could exhaust resources and trigger another kill, creating the bootloop.Warning
Firewall rules blocked me from connecting to one or more addresses (expand for details)
I tried to connect to the following addresses, but was blocked by firewall rules:
api.octopus.energy/home/REDACTED/work/batpred/batpred/coverage/venv/bin/python3 python3 ../apps/predbat/unit_test.py --quick(dns block)/home/REDACTED/work/batpred/batpred/coverage/venv/bin/python3 python3 ../apps/predbat/unit_test.py --quick git conf�� get --global rgo/bin/git HooksPath(dns block)/home/REDACTED/work/batpred/batpred/coverage/venv/bin/python3 python3 ../apps/predbat/unit_test.py --quick git conf�� get --global git core.hooksPath(dns block)gitlab.com/usr/lib/git-core/git-remote-https /usr/lib/git-core/git-remote-https origin REDACTED(dns block)https://api.github.com/repos/springfall2008/batpred/contents/apps/predbat/home/REDACTED/work/batpred/batpred/coverage/venv/bin/python3 python3 ../apps/predbat/unit_test.py --quick(http block)/home/REDACTED/work/batpred/batpred/coverage/venv/bin/python3 python3 ../apps/predbat/unit_test.py --quick git conf�� get --global rgo/bin/git HooksPath(http block)/home/REDACTED/work/batpred/batpred/coverage/venv/bin/python3 python3 ../apps/predbat/unit_test.py --quick git conf�� get --global git core.hooksPath(http block)If you need me to access, download, or install something from one of these locations, you can either: