Skip to content

fix: prevent a bare CancelledError during MCP connection from terminating the gateway #231

Description

@babyTa999

Summary

A failed streamable HTTP MCP connection can terminate the entire Raven gateway when the MCP SDK/anyio surfaces a bare asyncio.CancelledError.

connect_mcp_servers() currently catches (Exception, BaseExceptionGroup), but asyncio.CancelledError directly inherits from BaseException. It is therefore not caught by this handler and can escape the per-server connection boundary, terminating the gateway instead of logging a failed MCP connection and continuing startup.

In my case, the failure was triggered by an MCP server receiving no valid runtime authentication. After ensuring the Authorization header was available to the gateway process, the same server connected successfully and registered 25 tools. This fixes the local trigger, but does not address the uncaught-cancellation behavior.

Steps to reproduce

  1. Configure a streamable HTTP MCP server in Raven.
  2. Start Raven with invalid, missing, or otherwise rejected authentication for that MCP endpoint.
  3. Let the server/SDK cancel the streamable HTTP request during MCP initialization.
  4. Observe that a bare asyncio.exceptions.CancelledError can escape connect_mcp_servers() and terminate the entire gateway.

Relevant connection code currently uses:

except (Exception, BaseExceptionGroup) as e:
    logger.error("MCP server '{}': failed to connect: {}", name, e)

A bare asyncio.CancelledError is neither an Exception nor a BaseExceptionGroup.

Expected behavior

A single MCP server failing to connect should not terminate the Raven gateway.

Raven should:

  • log the failed MCP server name;
  • preserve useful context about the underlying transport/HTTP failure where possible;
  • skip that MCP server and continue starting the gateway.

External cancellation of the gateway/task should still propagate normally.

Actual behavior

The gateway terminated with a bare cancellation error similar to:

asyncio.exceptions.CancelledError: Cancelled via cancel scope ...

The original HTTP failure/status was not logged, so the underlying authentication rejection was not visible in the gateway log.

After valid runtime authentication was supplied, the MCP connection succeeded:

MCP server 'tanka': connected, 25 tools registered

Environment

  • OS: macOS arm64
  • Python: 3.12.13
  • Raven CLI: v0.1.3
  • MCP SDK: 1.28.1
  • httpx: 0.28.1
  • anyio: 4.14.2
  • MCP transport: streamable HTTP

Logs or screenshots

Please redact all credentials.

Failure:

asyncio.exceptions.CancelledError: Cancelled via cancel scope ...

Successful connection after valid runtime authentication:

MCP server 'tanka': connected, 25 tools registered

Activity

  1. 0xKT commented on Jul 29, 2026

    @0xKT
    Member

    Confirmed, and it is not specific to streamable HTTP or to an auth rejection -- we reproduced the same escape on the stdio transport, with no auth layer involved at all.

    Our CI hit it during a full unit-suite run (job log), on a commit that only touched a TypeScript comment:

    tests/test_sandbox_unit.py: in test_stdio_no_executor_does_not_raise
        await connect_mcp_servers({"svc": cfg}, ToolRegistry(), AsyncExitStack(), executor=None)
    raven/agent/tools/mcp.py: in connect_mcp_servers
        await session.initialize()
    mcp/client/session.py:171: in initialize
    asyncio/locks.py:212: CancelledError
    E   asyncio.exceptions.CancelledError: Cancelled via cancel scope 7fc5540452b0
    

    The cancellation escapes from await session.initialize(), which sits inside the same try your quoted handler belongs to:

    try:
    from mcp import ClientSession, StdioServerParameters
    from mcp.client.sse import sse_client
    from mcp.client.stdio import stdio_client
    from mcp.client.streamable_http import streamable_http_client
    if transport_type == "stdio":
    if executor is not None and executor.supports_process_spawning:
    read, write = await executor.start_process(cfg.command, cfg.args, env=cfg.env or None)
    else:
    params = StdioServerParameters(command=cfg.command, args=cfg.args, env=cfg.env or None)
    read, write = await stack.enter_async_context(stdio_client(params))
    elif transport_type == "sse":
    def httpx_client_factory(
    headers: dict[str, str] | None = None,
    timeout: httpx.Timeout | None = None,
    auth: httpx.Auth | None = None,
    ) -> httpx.AsyncClient:
    merged_headers = {**(cfg.headers or {}), **(headers or {})}
    return httpx.AsyncClient(
    headers=merged_headers or None,
    follow_redirects=True,
    timeout=timeout,
    auth=auth,
    )
    read, write = await stack.enter_async_context(
    sse_client(cfg.url, httpx_client_factory=httpx_client_factory)
    )
    elif transport_type == "streamableHttp":
    http_client = await stack.enter_async_context(
    httpx.AsyncClient(
    headers=cfg.headers or None,
    follow_redirects=True,
    timeout=None,
    )
    )
    read, write, _ = await stack.enter_async_context(
    streamable_http_client(cfg.url, http_client=http_client)
    )
    else:
    logger.warning("MCP server '{}': unknown transport type '{}'", name, transport_type)
    continue
    session = await stack.enter_async_context(ClientSession(read, write))
    await session.initialize()
    tools = await session.list_tools()
    for tool_def in tools.tools:
    wrapper = MCPToolWrapper(session, name, tool_def, tool_timeout=cfg.tool_timeout)
    registry.register(wrapper)
    logger.debug("MCP: registered tool '{}' from server '{}'", wrapper.name, name)
    logger.info("MCP server '{}': connected, {} tools registered", name, len(tools.tools))
    except (Exception, BaseExceptionGroup) as e:
    # BaseExceptionGroup is raised by anyio task groups (e.g. streamableHttp cancel
    # scope failures) and is not a subclass of Exception in Python 3.11+.
    logger.error("MCP server '{}': failed to connect: {}", name, e)

    Both transports reach initialize() through that one block, so a single fix there covers stdio and streamable HTTP together. The trigger on our side was just a stdio server whose process exits immediately -- a test configured with command="true" and spawning it for real. That test is being made hermetic in #230, so please do not rely on it as a reproducer once that merges.

    On the shape of the fix: the constraint that makes this more than widening the except is the one you already named -- external cancellation must still propagate. Catching BaseException, or CancelledError unconditionally, would swallow a genuine shutdown of the gateway's own task, so a fix has to distinguish "the MCP session's scope cancelled my frame" from "someone cancelled me". We have not settled on an implementation and are not asking you to; flagging it so the distinction is a deliberate decision rather than something a first patch runs into.

    One adjacent thing, separate from the except clause. The same CI run also produced:

    RuntimeError: Attempted to exit cancel scope in a different task than it was entered in
    

    In our run that came from the garbage collector finalizing a stdio_client async generator whose AsyncExitStack was never exited -- our test's fault, not the library's, so it is not evidence about your report. It is worth mentioning only because the production shape is structurally the same: the MCP stack is entered in AgentLoop._connect_mcp and closed later in close_mcp, which already swallows exactly this error:

    Raven/raven/agent/loop/main.py

    Lines 1866 to 1875 in 0640a25

    async def close_mcp(self) -> None:
    """Close MCP connections and the sandbox executor."""
    if self._mcp_stack:
    try:
    await self._mcp_stack.aclose()
    except (RuntimeError, BaseExceptionGroup):
    pass # MCP SDK cancel scope cleanup is noisy but harmless
    self._mcp_stack = None
    self._mcp_connected = False # reset so _connect_mcp() can reconnect after close
    self._mcp_connecting = False # reset so a concurrent caller isn't permanently blocked

    If a fix only widens the handler, that cancel-scope ownership question stays open, along with whatever the swallowed cleanup was failing to release.

  2. 0xKT commented on Jul 29, 2026

    @0xKT
    Member

    Two things found while reviewing #230 that bear on this: the codebase already
    defends against this exact SDK behaviour one function away, and the failure has
    a 3-line deterministic repro.

    The same file already guards for it on the call path

    raven/agent/tools/mcp.py:52 -- MCPToolWrapper.call:

    except asyncio.CancelledError:
        # MCP SDK's anyio cancel scopes can leak CancelledError on timeout/failure.
        # Re-raise only if our task was externally cancelled (e.g. /stop).
        task = asyncio.current_task()
        if task is not None and task.cancelling() > 0:
            raise
        logger.warning("MCP tool '{}' was cancelled by server/SDK", self._name)
        return "(MCP tool call was cancelled)"

    raven/agent/tools/mcp.py:167 -- connect_mcp_servers, the boundary this issue
    is about:

    except (Exception, BaseExceptionGroup) as e:
        # BaseExceptionGroup is raised by anyio task groups (e.g. streamableHttp cancel
        # scope failures) and is not a subclass of Exception in Python 3.11+.

    So the call path knows the SDK leaks a bare CancelledError and distinguishes
    "the SDK cancelled us" from "we were cancelled externally" via
    task.cancelling(). The connect path never got the same treatment. Whatever fix
    lands here, the task.cancelling() discrimination is the part worth carrying
    over -- swallowing every CancelledError at line 167 would make a real /stop
    during startup look like a failed server instead of a cancellation.

    This already fired in CI, from the test suite

    The unit run on #230 failed once with exactly this escape:

    FAILED tests/test_sandbox_unit.py::TestConnectMcpSandboxGuard::test_stdio_no_executor_does_not_raise
           - asyncio.exceptions.CancelledError: Cancelled via cancel scope 7fc5540452b0
    

    That test set command="true" and let connect_mcp_servers spawn it for real, so
    a child process racing the MCP handshake occasionally produced the leak this issue
    describes. The same test failed identically on main a day earlier
    (run 30338620584, macos-latest). It was never coverage -- it surfaced as a red
    build, not as an assertion -- and #230 has since made that test inject a raising
    transport instead of spawning, which is right for a unit test but does remove the
    accidental signal. Nothing in the suite exercises a CancelledError crossing the
    connect boundary today.

    Deterministic repro

    Using the injection idiom #230 just established for that test:

    import asyncio
    from contextlib import AsyncExitStack
    from unittest.mock import MagicMock
    
    import mcp.client.stdio
    
    from raven.agent.tools.mcp import connect_mcp_servers
    from raven.agent.tools.registry import ToolRegistry
    
    
    async def test_cancelled_error_does_not_escape_connect(monkeypatch):
        def fake_stdio_client(params):
            raise asyncio.CancelledError()
    
        monkeypatch.setattr(mcp.client.stdio, "stdio_client", fake_stdio_client)
    
        cfg = MagicMock()
        cfg.type = "stdio"
        cfg.command = "mcp-server"
        cfg.args = []
        cfg.env = None
        cfg.tool_timeout = 30
    
        # Today this raises CancelledError out of connect_mcp_servers.
        await connect_mcp_servers({"svc": cfg}, ToolRegistry(), AsyncExitStack(), executor=None)

    from mcp.client.stdio import stdio_client at mcp.py:116 is a deferred import
    inside the function, so patching the module attribute takes effect at call time.
    Swap stdio_client for streamable_http_client to reproduce the transport this
    issue was actually hit on.

  3. 0xKT commented on Oct 9, 2026

    @0xKT
    Member

    Fixed on main and included since v0.2.0. Each MCP server's transport, session and handshake now run as one lifecycle (#365, ported to raven/mcp/client.py in 7052056), so an anyio task-group failure becomes that server's connect error with the real cause logged, and MCPConnectionManager._handshake_watchdog turns any SDK-internal CancelledError into a connect failure while an external cancel still propagates (4e3c11c). Re-checked with MCP SDK 1.30.0 against a server answering 401: main logs "MCP server 'tanka_bad': failed to connect: HTTPStatusError: Client error '401 Unauthorized' ...", marks it auth_required and still connects the other servers, where the issue-time code (0640a25) raised CancelledError('Cancelled via cancel scope ...'); tests/test_mcp_manager.py::test_a_failed_transport_does_not_cancel_the_following_server guards this path. If the gateway still exits this way on a current build, please open a new issue with the log.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions