Skip to content

Fix dead retry-on-connection-failure logic in ArangoDatabase.connect() - #1331

Merged
tomchop merged 2 commits into
mainfrom
fix/connect-retry-exception-type
Jul 30, 2026
Merged

tomchop merged 2 commits into
mainfrom
fix/connect-retry-exception-type

Conversation

@tomchop

@tomchop tomchop commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator

Summary

Follow-up from investigating yeti-docker#29 (empty Feeds page on fresh
install) — a more foundational bug I found underneath the task-registration
race, affecting ArangoDatabase.connect() itself rather than just feed
bootstrap.

connect()'s startup retry loop (core/database_arango.py) is meant to
tolerate ArangoDB being briefly unreachable: retry the initial
has_database() check up to 4 times over ~20s, then log clearly and
sys.exit(1) rather than crash. It catches requests.exceptions.ConnectionError —
but python-arango 8.1.x (the pinned version) doesn't let that exception
escape. Its own host resolver
(arango/connection.py:178) catches the underlying transport error
internally and re-raises the builtin ConnectionError
(ConnectionAbortedError, specifically) instead. I verified
requests.exceptions.ConnectionError and the builtin ConnectionError
are unrelated sibling classes — both subclass OSError independently,
neither is a subclass of the other (issubclass(ConnectionAbortedError, requests.exceptions.ConnectionError) → False) — so the existing except
clause could never match what python-arango actually raises.

Net effect: any Yeti process (webserver, celery worker, celery beat,
events consumer) started before ArangoDB is ready to accept connections
crashes immediately with an unhandled ConnectionAbortedError traceback
on the very first connection attempt, instead of retrying for ~20s and
exiting cleanly. Worse, only the arangodb service has a restart
policy in the yeti-docker compose files — api/tasks/events-tasks/
tasks-beat don't — so a crash here doesn't even self-heal via Docker's
own restart mechanism.

This is a real, independently-reproducible bug regardless of the
compose-level healthcheck mitigation already shipped
(yeti-platform/yeti-docker#30) — it's relevant to any deployment without
equivalent readiness gating (bare docker run, Kubernetes without a
readiness probe), and to transient ArangoDB restarts after initial
startup (the yeti-docker history has at least one commit restarting
arangodb on OOM).

Fix

Catch both exception types — which one surfaces depends on where in the
request pipeline the failure originates, so there's no reason to pick
only one.

Test plan

  • Added test_connect_retries_and_exits_cleanly_on_unreachable_host
    (tests/core_tests/database_arango.py): points connect() at an
    unreachable host (mocking time.sleep to keep the test fast) and
    asserts it raises SystemExit (the clean exit(1) path) rather
    than letting ConnectionAbortedError escape unhandled.
  • Verified it fails with an uncaught ConnectionAbortedError on the
    pre-fix code, and passes with the fix.
  • Full tests/core_tests suite: 30/30 pass.
  • tests/schemas (190/190) and tests/apiv2 (198/200, same 2 known
    pre-existing failures) pass unchanged.
  • ty check (core+yetictl and plugins jobs): 0 errors.
  • ruff check / ruff format --check: clean.

connect()'s startup retry loop is meant to tolerate ArangoDB being
briefly unreachable: retry up to 4 times over ~20s, then log clearly
and exit(1) rather than crash. It caught
requests.exceptions.ConnectionError, but python-arango 8.1.x's own
host resolver (connection.py) catches the underlying transport error
internally and re-raises the builtin ConnectionError instead --
verified requests.exceptions.ConnectionError and the builtin
ConnectionError are unrelated sibling classes (both subclass OSError
independently), so the except clause could never match what
python-arango actually raises.

Net effect: any process started before ArangoDB is ready to accept
connections crashes immediately with an unhandled ConnectionAbortedError
traceback on the very first connection attempt, instead of retrying and
exiting cleanly. Only the arangodb service has a restart policy in the
yeti-docker compose files (api/tasks/events-tasks/tasks-beat don't), so
a crash here doesn't self-heal via Docker either.

Fix: catch both exception types, since which one surfaces depends on
where in the request pipeline the failure originates.

Added a regression test that points connect() at an unreachable host
(mocking time.sleep to keep it fast) and asserts it raises SystemExit
(the clean exit(1) path) rather than letting ConnectionAbortedError
escape unhandled -- verified it fails with an uncaught
ConnectionAbortedError on the pre-fix code and passes with the fix.
@tomchop
tomchop merged commit f85f0ff into main Jul 30, 2026
5 checks passed
@tomchop
tomchop deleted the fix/connect-retry-exception-type branch July 30, 2026 12:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant