Repository navigation
fix: prevent a bare CancelledError during MCP connection from terminating the gateway #231
Description
Activity
Confirmed, and it is not specific to streamable HTTP or to an auth rejection -- we reproduced the same escape on the stdio transport, with no auth layer involved at all.
Our CI hit it during a full unit-suite run (job log), on a commit that only touched a TypeScript comment:
tests/test_sandbox_unit.py: in test_stdio_no_executor_does_not_raise await connect_mcp_servers({"svc": cfg}, ToolRegistry(), AsyncExitStack(), executor=None) raven/agent/tools/mcp.py: in connect_mcp_servers await session.initialize() mcp/client/session.py:171: in initialize asyncio/locks.py:212: CancelledError E asyncio.exceptions.CancelledError: Cancelled via cancel scope 7fc5540452b0The cancellation escapes from
await session.initialize(), which sits inside the sametryyour quoted handler belongs to:Raven/raven/agent/tools/mcp.py
Lines 113 to 171 in 0640a25
try: from mcp import ClientSession, StdioServerParameters from mcp.client.sse import sse_client from mcp.client.stdio import stdio_client from mcp.client.streamable_http import streamable_http_client if transport_type == "stdio": if executor is not None and executor.supports_process_spawning: read, write = await executor.start_process(cfg.command, cfg.args, env=cfg.env or None) else: params = StdioServerParameters(command=cfg.command, args=cfg.args, env=cfg.env or None) read, write = await stack.enter_async_context(stdio_client(params)) elif transport_type == "sse": def httpx_client_factory( headers: dict[str, str] | None = None, timeout: httpx.Timeout | None = None, auth: httpx.Auth | None = None, ) -> httpx.AsyncClient: merged_headers = {**(cfg.headers or {}), **(headers or {})} return httpx.AsyncClient( headers=merged_headers or None, follow_redirects=True, timeout=timeout, auth=auth, ) read, write = await stack.enter_async_context( sse_client(cfg.url, httpx_client_factory=httpx_client_factory) ) elif transport_type == "streamableHttp": http_client = await stack.enter_async_context( httpx.AsyncClient( headers=cfg.headers or None, follow_redirects=True, timeout=None, ) ) read, write, _ = await stack.enter_async_context( streamable_http_client(cfg.url, http_client=http_client) ) else: logger.warning("MCP server '{}': unknown transport type '{}'", name, transport_type) continue session = await stack.enter_async_context(ClientSession(read, write)) await session.initialize() tools = await session.list_tools() for tool_def in tools.tools: wrapper = MCPToolWrapper(session, name, tool_def, tool_timeout=cfg.tool_timeout) registry.register(wrapper) logger.debug("MCP: registered tool '{}' from server '{}'", wrapper.name, name) logger.info("MCP server '{}': connected, {} tools registered", name, len(tools.tools)) except (Exception, BaseExceptionGroup) as e: # BaseExceptionGroup is raised by anyio task groups (e.g. streamableHttp cancel # scope failures) and is not a subclass of Exception in Python 3.11+. logger.error("MCP server '{}': failed to connect: {}", name, e) Both transports reach
initialize()through that one block, so a single fix there covers stdio and streamable HTTP together. The trigger on our side was just a stdio server whose process exits immediately -- a test configured withcommand="true"and spawning it for real. That test is being made hermetic in #230, so please do not rely on it as a reproducer once that merges.On the shape of the fix: the constraint that makes this more than widening the
exceptis the one you already named -- external cancellation must still propagate. CatchingBaseException, orCancelledErrorunconditionally, would swallow a genuine shutdown of the gateway's own task, so a fix has to distinguish "the MCP session's scope cancelled my frame" from "someone cancelled me". We have not settled on an implementation and are not asking you to; flagging it so the distinction is a deliberate decision rather than something a first patch runs into.One adjacent thing, separate from the
exceptclause. The same CI run also produced:RuntimeError: Attempted to exit cancel scope in a different task than it was entered inIn our run that came from the garbage collector finalizing a
stdio_clientasync generator whoseAsyncExitStackwas never exited -- our test's fault, not the library's, so it is not evidence about your report. It is worth mentioning only because the production shape is structurally the same: the MCP stack is entered inAgentLoop._connect_mcpand closed later inclose_mcp, which already swallows exactly this error:Raven/raven/agent/loop/main.py
Lines 1866 to 1875 in 0640a25
async def close_mcp(self) -> None: """Close MCP connections and the sandbox executor.""" if self._mcp_stack: try: await self._mcp_stack.aclose() except (RuntimeError, BaseExceptionGroup): pass # MCP SDK cancel scope cleanup is noisy but harmless self._mcp_stack = None self._mcp_connected = False # reset so _connect_mcp() can reconnect after close self._mcp_connecting = False # reset so a concurrent caller isn't permanently blocked If a fix only widens the handler, that cancel-scope ownership question stays open, along with whatever the swallowed cleanup was failing to release.
Two things found while reviewing #230 that bear on this: the codebase already
defends against this exact SDK behaviour one function away, and the failure has
a 3-line deterministic repro.The same file already guards for it on the call path
raven/agent/tools/mcp.py:52--MCPToolWrapper.call:except asyncio.CancelledError: # MCP SDK's anyio cancel scopes can leak CancelledError on timeout/failure. # Re-raise only if our task was externally cancelled (e.g. /stop). task = asyncio.current_task() if task is not None and task.cancelling() > 0: raise logger.warning("MCP tool '{}' was cancelled by server/SDK", self._name) return "(MCP tool call was cancelled)"
raven/agent/tools/mcp.py:167--connect_mcp_servers, the boundary this issue
is about:except (Exception, BaseExceptionGroup) as e: # BaseExceptionGroup is raised by anyio task groups (e.g. streamableHttp cancel # scope failures) and is not a subclass of Exception in Python 3.11+.
So the call path knows the SDK leaks a bare
CancelledErrorand distinguishes
"the SDK cancelled us" from "we were cancelled externally" via
task.cancelling(). The connect path never got the same treatment. Whatever fix
lands here, thetask.cancelling()discrimination is the part worth carrying
over -- swallowing everyCancelledErrorat line 167 would make a real/stop
during startup look like a failed server instead of a cancellation.This already fired in CI, from the test suite
The unit run on #230 failed once with exactly this escape:
FAILED tests/test_sandbox_unit.py::TestConnectMcpSandboxGuard::test_stdio_no_executor_does_not_raise - asyncio.exceptions.CancelledError: Cancelled via cancel scope 7fc5540452b0That test set
command="true"and letconnect_mcp_serversspawn it for real, so
a child process racing the MCP handshake occasionally produced the leak this issue
describes. The same test failed identically onmaina day earlier
(run 30338620584, macos-latest). It was never coverage -- it surfaced as a red
build, not as an assertion -- and #230 has since made that test inject a raising
transport instead of spawning, which is right for a unit test but does remove the
accidental signal. Nothing in the suite exercises aCancelledErrorcrossing the
connect boundary today.Deterministic repro
Using the injection idiom #230 just established for that test:
import asyncio from contextlib import AsyncExitStack from unittest.mock import MagicMock import mcp.client.stdio from raven.agent.tools.mcp import connect_mcp_servers from raven.agent.tools.registry import ToolRegistry async def test_cancelled_error_does_not_escape_connect(monkeypatch): def fake_stdio_client(params): raise asyncio.CancelledError() monkeypatch.setattr(mcp.client.stdio, "stdio_client", fake_stdio_client) cfg = MagicMock() cfg.type = "stdio" cfg.command = "mcp-server" cfg.args = [] cfg.env = None cfg.tool_timeout = 30 # Today this raises CancelledError out of connect_mcp_servers. await connect_mcp_servers({"svc": cfg}, ToolRegistry(), AsyncExitStack(), executor=None)
from mcp.client.stdio import stdio_clientatmcp.py:116is a deferred import
inside the function, so patching the module attribute takes effect at call time.
Swapstdio_clientforstreamable_http_clientto reproduce the transport this
issue was actually hit on.Fixed on main and included since v0.2.0. Each MCP server's transport, session and handshake now run as one lifecycle (#365, ported to raven/mcp/client.py in 7052056), so an anyio task-group failure becomes that server's connect error with the real cause logged, and MCPConnectionManager._handshake_watchdog turns any SDK-internal CancelledError into a connect failure while an external cancel still propagates (4e3c11c). Re-checked with MCP SDK 1.30.0 against a server answering 401: main logs "MCP server 'tanka_bad': failed to connect: HTTPStatusError: Client error '401 Unauthorized' ...", marks it auth_required and still connects the other servers, where the issue-time code (0640a25) raised CancelledError('Cancelled via cancel scope ...'); tests/test_mcp_manager.py::test_a_failed_transport_does_not_cancel_the_following_server guards this path. If the gateway still exits this way on a current build, please open a new issue with the log.
Summary
A failed streamable HTTP MCP connection can terminate the entire Raven gateway when the MCP SDK/anyio surfaces a bare
asyncio.CancelledError.connect_mcp_servers()currently catches(Exception, BaseExceptionGroup), butasyncio.CancelledErrordirectly inherits fromBaseException. It is therefore not caught by this handler and can escape the per-server connection boundary, terminating the gateway instead of logging a failed MCP connection and continuing startup.In my case, the failure was triggered by an MCP server receiving no valid runtime authentication. After ensuring the Authorization header was available to the gateway process, the same server connected successfully and registered 25 tools. This fixes the local trigger, but does not address the uncaught-cancellation behavior.
Steps to reproduce
asyncio.exceptions.CancelledErrorcan escapeconnect_mcp_servers()and terminate the entire gateway.Relevant connection code currently uses:
A bare
asyncio.CancelledErroris neither anExceptionnor aBaseExceptionGroup.Expected behavior
A single MCP server failing to connect should not terminate the Raven gateway.
Raven should:
External cancellation of the gateway/task should still propagate normally.
Actual behavior
The gateway terminated with a bare cancellation error similar to:
The original HTTP failure/status was not logged, so the underlying authentication rejection was not visible in the gateway log.
After valid runtime authentication was supplied, the MCP connection succeeded:
Environment
Logs or screenshots
Please redact all credentials.
Failure:
Successful connection after valid runtime authentication: