fix(integrations): preserve busy ACP sessions and defer rejected turns - #731
bandzalkin wants to merge 3 commits into
Conversation
bandzalkin
left a comment
There was a problem hiding this comment.
Architectural review → fix round
Scope: Reviewed #731 from merge base 8921c8ce8e5b78d4264a5e286880df87c74dc8a9 through correction commit ef1964426990d55474dc1276fcbaae5390326897, including protocol, ACP adapter, execution, regression, documentation, and changelog changes. Outcome: structured busy rejection applies backpressure without terminating the borrowed autonomous session or exhausting the delivery.
Constraints: Only RequestError(-32003, data.reason=session_busy) proves non-acceptance. Unknown or accepted work keeps ordinary timeout/cancellation semantics. ACP owns that proof; ExecutionContext owns watchdog, delivery retry, stop/play, and shutdown. Local contact events have no REST row to recover. Bootstrap/restoration provenance is removed with session cleanup.
Decisions: Keep a single total adapter deadline, fresh collector/emitter per request attempt, and a typed CancelledError carrying proven non-acceptance. The runtime converts only its own watchdog cancellation into deferral and preserves shutdown/control priority. Retain deferred local events at their FIFO position instead of inventing durable rows or consuming them. This is smaller than duplicating ACP error recognition in the runtime or creating another retry queue.
Findings and proof: A1 rejected-attempt contamination, A2 watchdog budget loss, and local-event loss were confirmed and corrected; cleanup provenance and shutdown races were checked. No standalone blocker remains. Fresh gate: 6,315 passed / 684 skipped; focused ACP/control tests 80 passed; all pre-commit hooks and both lock checks passed; baseline E2E collection 526 items / 1 skipped; markdown snippets 23 passed. Standalone real ACP socketpair smoke demonstrated watchdog deferral, same-session recovery, and an accepted retry's visible answer after unrelated autonomous tool output.
Architectural verdict: APPROVED for #731 independently. Generic OMP 18.6.0 errors remain ordinary errors; structured 18.6.1 rejection is supported. Combined integration with #729 is HELD: its live external-effect recorder commits rejected-attempt effects to the shared ledger before prompt success. The minimal staging correction belongs in #729, not a second policy in #731. This is not remote CI or merge approval.
Summary
Treat a proven ACP busy rejection as backpressure, not a provider failure that tears down the session. Preserve the autonomous turn and retry the room prompt against the same runtime/session.
This is SDK-only. No upstream OMP changes or requests are included.
Changes
acp.RequestErrorwith code-32003and dictionarydata.reason == "session_busy", using asynchronous exponential backoff (0.25 seconds, capped at 5 seconds) inside the existing turn timeout.TurnDeferred. Both backlog and WebSocket execution leave the delivery failed/actionable without spending its ordinary retry budget, acknowledging it as processed, or posting a terminal failure.session/cancelwhile waiting after a busy rejection. Preserve existing cancellation/timeout behavior once a prompt might have started; do not replay transport errors or unknown-acceptance failures through this path.OMP 18.6.1 emits the structured rejection. OMP 18.6.0's generic
-32603response remains an ordinary failure; message-text matching is deliberately not supported. This does not claim ownership of autonomous OMP turns or guarantee forwarding their unsolicited output.Related Issues
No issue linked. Independent companion SDK PR: #730 (backlog recovery past retained exhausted messages).
Testing
CI=true uv run --no-sync python -m pytest -q --no-cov— 6,309 passed, 684 skipped, 51 warnings (macOS, Python 3.13.14).BUSY_RECOVERY_OK.The smoke did not write to a live Band room. Opt-in live E2E/Docker lanes were not run. Remote CI results are separate from these local checks.
Checklist