Skip to content

MCP server returning 401 is retried forever (18+ failures, every 300 s) with no needs-auth state in the UI #6412

Description

@Al629176

Description

An MCP server whose credential was never saved (https://api.inference.sh, HTTP 401) is retried forever. Backoff grows to 300 s and then stays there; 18+ failures logged in one session with no circuit breaker, no UI indicator that the server needs auth, and no way to see it from the MCP Servers page.

Reported by: Alan (QA)
Build: v0.63.29
Reproducible: Yes

Steps to reproduce

  1. Add an MCP server that requires a bearer token; do not supply the token (or supply a wrong one)
  2. Observe:
WRN reconnecting failed: mcp unauthorized for `https://api.inference.sh` (HTTP 401) server_id=7aa4062b-… failures=1 retry_in_seconds=5
   … failures=2 → 10s, 3 → 20s, 4 → 40s, 5 → 80s, 6 → 160s, 7 → 300s … failures=18 retry_in_seconds=300 (continuing)
INF [jsonrpc] unconfirmed unauthorized error for method='openhuman.mcp_clients_connect' (not session expiry) — leaving session intact

Expected behaviour

  • A 401 is terminal until the credential changes: stop retrying, mark the server "needs authentication" in the MCP Servers list, and offer a fix action
  • Retry resumes only when the user updates the credential

Actual behaviour

  • Infinite retry every 5 min for the life of the session; nothing visible in the UI

Impact

  • Log/Sentry noise; wasted requests; user has no idea the server is dead
  • Combined with the secret-prompt bug, an auth-requiring MCP server can never reach a working state

Acceptance criteria

  • 401/403 from an MCP server halts reconnection and sets a needs_auth state visible on the MCP Servers page
  • A "Set token" action re-enables connection
  • No more than one 401 warning per server per session
  • Reproducible on v0.63.29 — confirm fix in next build

Activity

  1. added
    mcp-rpcMCP transport, tool registry, JSON-RPC, and core relay surfaces.
    priority: p2Soon. Real but survivable — a rough edge, a gap, a thing that will bite later.
    on Sep 21, 2026
  2. added
    priority: p1Next. Wrong behaviour a user will hit, or a security weakness behind a condition.
    and removed
    priority: p2Soon. Real but survivable — a rough edge, a gap, a thing that will bite later.
    on Sep 21, 2026
  3. tinysweeper commented on Sep 21, 2026

    @tinysweeper

    MCP server returning 401 is retried forever with exponential backoff capped at 300s, causing log noise and no user-visible auth-required state. Fix should halt retries on 401/403, mark the server as needing auth in the MCP Servers page, and provide a Set token action.

    Labelled priority: p1.

  4. added theissue type on Sep 21, 2026
  5. M3gA-Mind commented on Sep 22, 2026

    @M3gA-Mind
    Collaborator

    Still reproduces on 0f1ecc9d28d85289ee3f87ade13156be02ffbd28 — and the reason is a delivery gap, not a missing fix.

    The fix exists upstream. tinyhumansai/tinymcp#20 ("Park a server whose credential was rejected instead of retrying it forever") merged 2026-09-22T20:48Z as 10786a472e3c17b796c4c173a0e4647ed916b120.

    It has not been delivered here. vendor/tinymcp on main is pinned at fe34f5b8c89a49e6f7ca05c2da791bb34550f85f:

    $ git ls-tree 0f1ecc9d2 vendor/tinymcp
    160000 commit fe34f5b8c89a49e6f7ca05c2da791bb34550f85f	vendor/tinymcp
    
    $ gh api repos/tinyhumansai/tinymcp/compare/10786a472...fe34f5b8c --jq '{status,behind_by}'
    {"status":"behind","behind_by":2}
    

    The pinned commit is dated 2026-09-19; the fix merged three days later. The pin is behind the fix, so it cannot contain it. Every build cut from main still runs the old retry loop.

    What's needed: a vendor/tinymcp gitlink bump to a tinymcp commit at or after 10786a472. Nothing in this repo needs changing — the retry/backoff and the needs_auth state are both submodule-side.

    Test that would prove delivery: assert the pin contains the fix rather than eyeballing it, e.g. a CI step running git -C vendor/tinymcp merge-base --is-ancestor 10786a472 HEAD. More generally, the recurring failure mode here is that a merged submodule PR reads as "fixed" while users still run the old code — #6411 and #6415 are in the same position against the same pin.

  6. moved this from Todo to In progress in Team Openhumanon Sep 22, 2026
  7. added a commit that references this issue on Sep 23, 2026
  8. moved this from In progress to Done in Team Openhumanon Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugmcp-rpcMCP transport, tool registry, JSON-RPC, and core relay surfaces.priority: p1Next. Wrong behaviour a user will hit, or a security weakness behind a condition.

    Type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions