Repository navigation
fix: Replace a crashed Band MCP server in every adapter that hosts one - #727
Conversation
BandMCPBackend.restart() stops and starts the same LocalMCPServer; the server tries the port it last served on only after every other one in its range, so consumers holding the old URL can tell a restart happened. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
…essions Each message checks the backend under a lock and restarts it in place when its serve task died; cleanup_all detaches it under the same lock so nothing restarts after shutdown. ClaudeSessionManager replaces a cached client whose MCP servers no longer match the room's, resuming its session on the new URL, and fails requests queued behind stop() instead of leaving them hung. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
The crash branch restarts the backend in place, and every message with injected Band tools checks it. A room whose session was built against a URL the backend no longer serves gets a fresh session on the live one, with the transcript replayed, instead of failing until a turn tears it down. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
Every turn now checks the backend before the already-registered short-circuit; a dead one restarts in place on a new port and is re-registered under the same name, which OpenCode treats as a replacement. A failed restart propagates instead of letting the turn run without tools. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
The liveness check, warning and in-place restart lived in each adapter; BandMCPBackend now owns them and reports whether it restarted, which OpenCode uses to drop its stale registration. The OpenCode restart test waits for the first turn to finish before crashing the server. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
A self-hosted backend whose serve task died no longer counts as ready. The next message restarts it and updates the registration's URL in place: Letta reads that URL on every tool call, so attached tools reconnect with no new ids. A row still pointing at this process's own last URL is adopted rather than rejected, so a fixed server_name survives a crash after release. letta-client's floor rises to 1.0.0, the first release with the mcp_servers API the bridge already calls. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
start() only stopped what it had begun on Exception, so a cancellation mid-start left the serve task running and its port bound with no caller holding the server to stop it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
…MCPBackend claude_sdk, ACP, OpenCode and Letta each hand-rolled starting, healing, locking and shutting down their Band MCP server, so a fifth adapter could copy one and miss the crash check. SharedBandMCPBackend now owns all of it: ensure() starts or heals the backend and refuses after a final close, close()/detach() stop it outside the lock, and reopen() re-arms it on agent restart. Adapters declare BandMCPBackendSettings and keep only their own reconnect step, keyed on the backend's URL changing. A guard test fails any module outside band.integrations.mcp that creates, constructs or heals a Band MCP server itself. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
… place A restarted-in-place backend kept its identity while its port changed, so correctness leaned on ordering rules (read the endpoint before any await) that held only because nothing could run while the port was unset. The owner now stops a dead backend and starts a new one off its port, so a BandMCPBackend's URLs are fixed for its whole life and a stopped server still reports the URLs it served. restart_if_crashed() and LocalMCPServer's remembered previous port give way to an explicit avoid_port. ACP records each room's Band MCP URL on the RoomSession it belongs to, written once the session exists and retired with it, replacing a parallel dict kept in step at three sites. Also: OpenCode compares registrations by value, Letta builds its registration config and URL once, and tests share one fake backend plus hold/crash/served-tools helpers in tests/mcpclient.py. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
…registration ACP decides whether a room's session is still current inside _get_or_create_session, from one Band URL it also builds the session with, instead of a separate stale check on_message had to run first. Letta's ready compares the registered URL with the live server's, so a repoint interrupted mid-update is retried rather than trusted; its repoint and forget-registration steps each live in one place. Also: stale docstrings and the OpenCode shutdown comment now match the replace-don't-restart design, the restart test files and OpenCode's backend fixture are renamed for it, and two tests that never dial their server run on a fake. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
AlexanderZ-Band
left a comment
There was a problem hiding this comment.
Review of the replace-on-crash design at a030cea. Findings are inline, plus one for a line outside the diff:
tests/mcpclient.py:1— suggestion: the module docstring says "Real Band MCP servers and client sessions for tests that dial them", but the module now also holds the backend fakes (FakeLocalServer,FakeBandMCPBackend,BackendStarts,backends_created_by,hold_backend), which exist for tests that never dial a server. Two jobs sit in one module, and its docstring describes only one of them.
Blocking: two confirmed defects in new release paths (session_manager.py:144 restarts a stopped manager; client_adapter.py:1010 leaks Cursor todo state on replacement), plus test-quality gaps.
…rsor sessions A ClaudeSessionManager stop is now final: the adapter builds a new manager to start again, so a session requested after stop() raises ClaudeSessionManagerStoppedError, which on_message passes through instead of retrying as a failed resume that restarted the stopped loop. ACP releases a session through one _release_session hook, from room cleanup and from replacing a session whose Band MCP URL went stale, so CursorACPAdapter forgets the retired session's todo state either way. Also from review: ACP builds a session's MCP servers once per reuse decision; backend fakes move to tests/mcpbackends.py; one shared SHIPPED_SOURCE_ROOTS for the AST guards; a replacement logs the dead port; tests assert observable outcomes, spec Letta's API from letta-client, make the concurrent-start test actually overlap, and drive the post-shutdown ACP turn through the harness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
AlexanderZ-Band
left a comment
There was a problem hiding this comment.
Cycle 2 of /my-code-review: 14 findings (F logic bugs ×3, E ×1, D ×3, C ×1, B ×2, A ×4). Fixes follow.
…r their runtime stops - ClaudeSessionManager.stop() runs one shielded shutdown that every caller awaits, so overlapping or cancelled stops no longer fail or abandon it; lifecycle state is the loop task and the shutdown task, not two flags. - ClaudeSDKAdapter's resume fallback lets a stopped manager pass through quietly instead of posting a failure mid-shutdown. - ACPClientAdapter.on_cleanup releases the session after runtime.stop(), so a detached turn's late Cursor todo updates can't outlive the cleanup. - The backend-ownership guard moves to its own module and also catches module-qualified calls; tests pin the drain, the retired session's runtime reset, and the fixes above. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
AlexanderZ-Band
left a comment
There was a problem hiding this comment.
Cycle 3 of /my-code-review: 9 findings. Two of them (E, C) share one root cause in cycle 2's ACP release reorder. The logic-bug explorer's H3 was dropped: it isn't reachable through Agent.stop, and when forced it posts nothing, whereas the base branch hangs.
…s state even on a cancelled stop - ACPClientAdapter drops a session's bootstrap mark under _session_lock with no await before it again (on_cleanup and the stale-session branch), and calls the new _forget_session hook in a finally once the session can deliver nothing more: after on_cleanup's runtime.stop() and after every _close_session. Cursor overrides _forget_session. - ClaudeSessionManager's shutdown queues a future-less stop and awaits the loop task; the stop future and _queue_stop are gone. - SharedBandMCPBackend refuses with BandMCPBackendStoppedError. - The ownership guard also catches an aliased LocalMCPServer import. - Tests: a cleanup cancelled mid-stop, a manager without an MCP factory reusing its session; the fake ACP agent can hold its exit and has one Cursor-todos sender. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
…r a reload of the retired one A bootstrap turn names the room's persisted session; when that session's Band MCP URL is stale, the initializer restored the very id being retired and the background close then killed it. The stale branch now drops the restore, so the room gets the fresh session the retire intends. Also pins the session manager's no-loop guards: cleanup, invalidate and cleanup_all return at once on a never-started or stopped manager. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN
AlexanderZ-Band
left a comment
There was a problem hiding this comment.
Cycle 4 reviewed the PR against transition tables for its two churn areas: the ACP session lifetime and the ClaudeSessionManager lifecycle. Every row holds. One missing row, a gap in the PR's own design, is filed inline. Two untested rows (commands on a never-started or a stopped manager) are now pinned.
AlexanderZ-Band
left a comment
There was a problem hiding this comment.
Follow-up filed for defects on the base branch: INT-1666.
amit-gazal-band
left a comment
There was a problem hiding this comment.
Approving — nice rework. SharedBandMCPBackend and replace-instead-of-restart resolve most of what I flagged on 627fd201: the restarted session manager, and endpoint() raising during a restart.
All inline comments are optional follow-ups (races and edge cases, nothing blocking). The PR is approved regardless.
Verified locally at 4723c38f: unit tests 6207 passed / 151 skipped / 1 failed (test_strands_adapter::test_turn_agent_runs_the_configured_model_id, a local AWS botocore[crt] problem, unrelated). tests/mcp/test_backend_ownership.py + test_import_boundary.py pass. ruff check/format clean. The pyrefly errors are only band_mcp imports that my dev extra doesn't install.
🤖 Generated with Claude Code
amit-gazal-band
left a comment
There was a problem hiding this comment.
Approving. Thanks for the rework: SharedBandMCPBackend and replace-instead-of-restart resolve the session-manager restart race and the endpoint()-mid-restart crashes from my earlier review.
I left 3 inline comments with small fixes I'd recommend, all low-effort.
Fine as follow-up PRs (not needed for this one):
- Same-port replacement isn't detected (
backends.py_replacement). Every consumer (Claude_is_current, ACPband_url, OpenCodeMcpRegistration) detects a replaced server by comparing URLs, butavoid_portonly tries the old port last. When the range is nearly full (or has one port), the replacement can bind the same port, and sessions keep transport state the new server never issued. A generation counter onSharedBandMCPBackend, compared by consumers, would detect this reliably and remove the per-adapter URL tracking. - ACP can return a just-retired session (
client_adapter.py_get_or_create_session). After a stale mapping is deleted, a finished, successful initializer that hasn't been released yet is still reused. It returns the session id being closed ascreated=False, with no transcript replay. The window is narrow (overlapping turns in one room plus a crash). Dropping finished initializers, or checking theirband_url, closes it.
Verified locally at 4723c38f: 6207 passed / 151 skipped. The one failure, test_strands_adapter::test_turn_agent_runs_the_configured_model_id, is a local AWS botocore[crt] problem and unrelated. The new tests/mcp ownership/import-boundary tests pass, and ruff check/format is clean.
🤖 Generated with Claude Code
…eep Letta's registration on a transient repoint failure, and drop a retired ACP session's finished setup - ClaudeSessionManager: a failing cleanup no longer keeps the loop alive past stop, and teardown commands queued behind stop resolve instead of raising; only a create fails. - Letta MCP bridge: a repoint failure forgets the registration only on a 404, so a transient error retries the same row instead of registering a second one. - ACP client adapter: a finished setup that built the session just retired is never handed to the next turn. Co-authored-by: Cursor <cursoragent@cursor.com>
45107fc
into
fix/fix-bind-the-room-server-side-so-models-never-supp-INT-1607
… chat_id (#723) * fix!: bind the room to the Band MCP connection so models never supply chat_id A room-bound LocalMCPServer mounts its engine under /rooms/{room_id}; its tools advertise no chat_id and each call takes its room from the request path. ACP client sessions and claude_sdk dial their room's endpoint via BandMCPBackend.endpoint(); opencode and letta keep the multi-room endpoint. claude_sdk moves off the in-process SDK MCP server to HTTP, and the chat_id prompt text is dropped wherever the tools are bound. BREAKING CHANGE: the "sdk" Band MCP backend kind is removed; use kind="http" (with room_bound=True for per-room endpoints). BandMCPBackend drops its `server` field (use `local_server`), create_band_mcp_backend drops get_participant_handles/tool_result_hook, and the Claude SDK tool builders in band.integrations.claude_sdk.tools are removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017VPm1kcnGmxdLivQ6BGJj5 * refactor: share one transport vocabulary and room-endpoint test helpers MCPTransportKind aliases BandMCPBackendKind; the ACP prompt builds its room block under one conditional; tests address room endpoints through room_endpoint_path and start backends through one started_backend helper. The copilot_sdk prompt test now asserts the prompt carries no room id. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017VPm1kcnGmxdLivQ6BGJj5 * docs: drop stale chat_id and in-process wording from claude_sdk docstrings Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017VPm1kcnGmxdLivQ6BGJj5 * refactor: remove the claude_sdk tool type aliases left unused Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017VPm1kcnGmxdLivQ6BGJj5 * refactor!: drop the unread Band MCP backend kind and name its transport type Every backend is one LocalMCPServer serving both transports, so `kind` was never read; consumers pick the transport in `endpoint()`. The one transport vocabulary is now `BandMCPTransport` (replacing `BandMCPBackendKind` and the ACP-only `MCPTransportKind`). Docs and examples describe the room-bound endpoints ACP and claude_sdk now use. BREAKING CHANGE: `create_band_mcp_backend` and `BandMCPBackend` no longer take or carry `kind`; `BandMCPBackendKind` is renamed `BandMCPTransport`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017VPm1kcnGmxdLivQ6BGJj5 * fix(mcp): serve Band MCP tool results as plain JSON text, not a structured wrapper Every engine tool returns its JSON as text, but FastMCP inferred a `{"result": string}` output schema from the handler's `-> str` and sent the same text again as `structuredContent: {"result": "<json>"}`. Clients that prefer structured content (the Claude CLI; Copilot, which the ACP echo unwrap exists for) showed the model double-encoded JSON. claude_sdk moved onto the engine in this PR, so its roster lookups started reaching the model escaped. The engine now builds every tool unstructured, which also drops the per-registration `structured_output` field it no longer needs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017VPm1kcnGmxdLivQ6BGJj5 * refactor(mcp): tighten the Band MCP backend's public surface - create_band_mcp_backend types get_tools as RoomToolResolver (was Any) and takes any Sequence of tool definitions. - BandMCPBackend is frozen; endpoint()'s room_id defaults to None. - build_custom_tool_registration takes one advertise_chat_id flag instead of a room_bound/room_from_connection pair that described one choice. - The connection-room reader is private to the engine. - ClaudeSessionManager's mcp_servers_factory returns the SDK's own McpServerConfig mapping (was dict[str, Any]). - LocalMCPServer's wrong-kind URL error names the accessor to use. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017VPm1kcnGmxdLivQ6BGJj5 * test(claude_sdk): give the turn-timeout tests a deadline a real MCP round trip fits The deadline is adapter-wide, so the turn after the held one must finish within it too. Its Band tool calls now cross a real loopback HTTP server, which a 50 ms deadline did not leave room for on CI runners. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017VPm1kcnGmxdLivQ6BGJj5 * test(e2e): report why an approval turn never closed A stalled approval turn timed out with a bare TimeoutError: the outer deadline always beat the inner one and swallowed its message, so a silent agent and a frame the WebSocket capture missed looked the same. The timeout now outlines, without message contents, each streamed and durable agent message's role (request, notice, closing reply, other text) and whether the durable copy closed. Phase logs mark the captured requests and the moment the stream settles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017VPm1kcnGmxdLivQ6BGJj5 * fix: restore MCP tool failure logs and guard chat_id-free prompts Cycle-1 review: log unexpected engine dispatch failures (FastMCP drops the stack), pin the claude_sdk prompt drift guard, and finish the mcp_servers_factory docstring. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: tighten room-bound kind-mismatch tests and clarify pin docs Cycle-2 review: pin each endpoint-kind assertion to its expected error fragment, and describe pin_existing_chat_id by what the schema does for both callers. Co-authored-by: Cursor <cursoragent@cursor.com> * refactor(mcp): make BandMCPTransport a StrEnum so each transport has one definition Callers spelled the transport as a bare "http"/"sse" literal at every site (opencode, claude_sdk, ACP's capability selection, the tests). They now reference BandMCPTransport.HTTP/.SSE, and the endpoint and ACP config dispatch match on the enum. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017VPm1kcnGmxdLivQ6BGJj5 * fix: read room-bound MCP rooms from FastMCP Context Stop importing mcp's low-level request_ctx ContextVar; inject ctx: Context into room-bound dispatch so upgrades can't silently break room binding. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: Replace a crashed Band MCP server in every adapter that hosts one (#727) * fix: Restart a Band MCP server in place on a new port BandMCPBackend.restart() stops and starts the same LocalMCPServer; the server tries the port it last served on only after every other one in its range, so consumers holding the old URL can tell a restart happened. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * fix: Restart claude_sdk's crashed Band MCP server and recycle stale sessions Each message checks the backend under a lock and restarts it in place when its serve task died; cleanup_all detaches it under the same lock so nothing restarts after shutdown. ClaudeSessionManager replaces a cached client whose MCP servers no longer match the room's, resuming its session on the new URL, and fails requests queued behind stop() instead of leaving them hung. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * fix: Heal ACP rooms whose session dials a crashed Band MCP server The crash branch restarts the backend in place, and every message with injected Band tools checks it. A room whose session was built against a URL the backend no longer serves gets a fresh session on the live one, with the transcript replayed, instead of failing until a turn tears it down. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * fix: Re-register OpenCode's Band MCP server after it crashes Every turn now checks the backend before the already-registered short-circuit; a dead one restarts in place on a new port and is re-registered under the same name, which OpenCode treats as a replacement. A failed restart propagates instead of letting the turn run without tools. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * refactor: Heal every Band MCP backend through one restart_if_crashed() The liveness check, warning and in-place restart lived in each adapter; BandMCPBackend now owns them and reports whether it restarted, which OpenCode uses to drop its stale registration. The OpenCode restart test waits for the first turn to finish before crashing the server. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * fix: Repoint Letta's registration at a restarted Band MCP server A self-hosted backend whose serve task died no longer counts as ready. The next message restarts it and updates the registration's URL in place: Letta reads that URL on every tool call, so attached tools reconnect with no new ids. A row still pointing at this process's own last URL is adopted rather than rejected, so a fixed server_name survives a crash after release. letta-client's floor rises to 1.0.0, the first release with the mcp_servers API the bridge already calls. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * fix: Clean up a cancelled LocalMCPServer start start() only stopped what it had begun on Exception, so a cancellation mid-start left the serve task running and its port bound with no caller holding the server to stop it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * refactor: Own every adapter's Band MCP backend through one SharedBandMCPBackend claude_sdk, ACP, OpenCode and Letta each hand-rolled starting, healing, locking and shutting down their Band MCP server, so a fifth adapter could copy one and miss the crash check. SharedBandMCPBackend now owns all of it: ensure() starts or heals the backend and refuses after a final close, close()/detach() stop it outside the lock, and reopen() re-arms it on agent restart. Adapters declare BandMCPBackendSettings and keep only their own reconnect step, keyed on the backend's URL changing. A guard test fails any module outside band.integrations.mcp that creates, constructs or heals a Band MCP server itself. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * refactor: Replace a dead Band MCP backend instead of restarting it in place A restarted-in-place backend kept its identity while its port changed, so correctness leaned on ordering rules (read the endpoint before any await) that held only because nothing could run while the port was unset. The owner now stops a dead backend and starts a new one off its port, so a BandMCPBackend's URLs are fixed for its whole life and a stopped server still reports the URLs it served. restart_if_crashed() and LocalMCPServer's remembered previous port give way to an explicit avoid_port. ACP records each room's Band MCP URL on the RoomSession it belongs to, written once the session exists and retired with it, replacing a parallel dict kept in step at three sites. Also: OpenCode compares registrations by value, Letta builds its registration config and URL once, and tests share one fake backend plus hold/crash/served-tools helpers in tests/mcpclient.py. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * refactor: Decide reuse where it happens for ACP sessions and Letta's registration ACP decides whether a room's session is still current inside _get_or_create_session, from one Band URL it also builds the session with, instead of a separate stale check on_message had to run first. Letta's ready compares the registered URL with the live server's, so a repoint interrupted mid-update is retried rather than trusted; its repoint and forget-registration steps each live in one place. Also: stale docstrings and the OpenCode shutdown comment now match the replace-don't-restart design, the restart test files and OpenCode's backend fixture are renamed for it, and two tests that never dial their server run on a fake. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * fix: Stop restarting a stopped session manager and leaking retired Cursor sessions A ClaudeSessionManager stop is now final: the adapter builds a new manager to start again, so a session requested after stop() raises ClaudeSessionManagerStoppedError, which on_message passes through instead of retrying as a failed resume that restarted the stopped loop. ACP releases a session through one _release_session hook, from room cleanup and from replacing a session whose Band MCP URL went stale, so CursorACPAdapter forgets the retired session's todo state either way. Also from review: ACP builds a session's MCP servers once per reuse decision; backend fakes move to tests/mcpbackends.py; one shared SHIPPED_SOURCE_ROOTS for the AST guards; a replacement logs the dead port; tests assert observable outcomes, spec Letta's API from letta-client, make the concurrent-start test actually overlap, and drive the post-shutdown ACP turn through the harness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * fix: Share one session-manager shutdown and release ACP sessions after their runtime stops - ClaudeSessionManager.stop() runs one shielded shutdown that every caller awaits, so overlapping or cancelled stops no longer fail or abandon it; lifecycle state is the loop task and the shutdown task, not two flags. - ClaudeSDKAdapter's resume fallback lets a stopped manager pass through quietly instead of posting a failure mid-shutdown. - ACPClientAdapter.on_cleanup releases the session after runtime.stop(), so a detached turn's late Cursor todo updates can't outlive the cleanup. - The backend-ownership guard moves to its own module and also catches module-qualified calls; tests pin the drain, the retired session's runtime reset, and the fixes above. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * fix: Forget an ACP session's bootstrap under the lock and its subclass state even on a cancelled stop - ACPClientAdapter drops a session's bootstrap mark under _session_lock with no await before it again (on_cleanup and the stale-session branch), and calls the new _forget_session hook in a finally once the session can deliver nothing more: after on_cleanup's runtime.stop() and after every _close_session. Cursor overrides _forget_session. - ClaudeSessionManager's shutdown queues a future-less stop and awaits the loop task; the stop future and _queue_stop are gone. - SharedBandMCPBackend refuses with BandMCPBackendStoppedError. - The ownership guard also catches an aliased LocalMCPServer import. - Tests: a cleanup cancelled mid-stop, a manager without an MCP factory reusing its session; the fake ACP agent can hold its exit and has one Cursor-todos sender. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * fix: Give an ACP room whose Band URL went stale a fresh session, never a reload of the retired one A bootstrap turn names the room's persisted session; when that session's Band MCP URL is stale, the initializer restored the very id being retired and the background close then killed it. The stale branch now drops the restore, so the room gets the fresh session the retire intends. Also pins the session manager's no-loop guards: cleanup, invalidate and cleanup_all return at once on a never-started or stopped manager. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN * fix: Settle teardown commands queued behind a session-manager stop, keep Letta's registration on a transient repoint failure, and drop a retired ACP session's finished setup - ClaudeSessionManager: a failing cleanup no longer keeps the loop alive past stop, and teardown commands queued behind stop resolve instead of raising; only a create fails. - Letta MCP bridge: a repoint failure forgets the registration only on a 404, so a transient error retries the same row instead of registering a second one. - ACP client adapter: a finished setup that built the session just retired is never handed to the next turn. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix: tear down Claude SDK adapters under looptime Fixture teardown runs on the real clock after looptime tests, so a uvicorn tick scheduled at virtual T blocked LocalMCPServer.stop() for ~T wall seconds on fresh CI runners. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: access looptime teardown helpers via getattr Keep AbstractEventLoop typing happy without casting the running loop. Co-authored-by: Cursor <cursoragent@cursor.com> * test: preserve the MCP server clock during Claude fixture cleanup --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Stacked on #723. Retarget to
mainonce it merges.Resolves INT-1642.
Why
claude_sdk, ACP, OpenCode and Letta each host a loopback Band MCP server and never check it again. When its serve task dies, every session or registration keeps dialing a dead port. The model loses its Band tools until the agent restarts. Each adapter also hand-rolled its own start, lock and shutdown logic, so a new adapter could easily miss the crash check.
What
SharedBandMCPBackend, used by all four adapters:ensure()starts the server, replaces it if it died, and refuses after a final shutdown.close()stops it;reopen()re-arms it when the agent restarts.BandMCPBackendSettings.band.integrations.mcpcreates a Band MCP server itself.LocalMCPServer.start()no longer leaks a running server.ClaudeSessionManager.stop()settles queued requests instead of leaving them hanging (teardown calls succeed, new-session requests fail), and a stopped manager refuses new sessions instead of quietly starting again.letta-clientfloor is now>=1.0.0; earlier releases lack themcp_serversAPI.Follow-up
opencode serve.🤖 Generated with Claude Code
https://claude.ai/code/session_013rPu5kStSAEAxctEv6Z9uN