Skip to content

Orchestrator's Temporal worker never retries a failed initial connection — silently stops polling forever, no crash, no log after the first attempt #2035

Description

@ptigroup

Orchestrator's Temporal worker never retries a failed initial connection — silently stops polling forever, no crash, no log after the first attempt

Versions: nestjs-temporal-core@3.2.3, @temporalio/worker@1.15.0 (self-hosted, docker-compose deployment)

What happened

After a docker compose restart that raced Temporal's own startup, the orchestrator app's Temporal worker failed its one connection attempt at boot:

[TemporalWorkerManagerService] ERROR Failed to create connection
TransportError: tcp connect error, <temporal-ip>:7233, Connection refused

The NestJS app then logged Nest application successfully started and kept running normally — HTTP health endpoint up, pm2/process manager showing it "online" with 0 restarts — but the Temporal worker never registered a poller on any task queue again. Scheduled posts (workflows already queued) sat untouched indefinitely, because nothing was ever polling for them. This condition lasted 31 hours before being noticed, and is a repeat of an earlier ~week-long occurrence of the same failure mode.

Root cause

initializeMultipleWorkers() (and the single-worker initializeWorker() path) in TemporalWorkerManagerService treats a connection failure at boot as tolerable by default (allowConnectionFailure !== false): it logs a warning/error and returns, rather than retrying or throwing. onModuleInit() runs exactly once, at process startup — there is no later code path (no interval, no scheduled retry, no reconnect-on-next-use) that ever attempts the connection again. A transient "Temporal isn't ready yet" condition at boot therefore becomes a permanent loss of worker functionality for the life of the process, indistinguishable from healthy at the process-supervisor level.

Why this is worse than a normal transient-connection issue

Most supervisors (pm2, Docker, k8s) recover automatically from a crashed process. This failure mode specifically avoids that safety net: the process doesn't crash, so nothing restarts it, and its own health endpoint doesn't reflect the missing worker either (in Postiz's case, /health/status opens an independent fresh Temporal connection to test Temporal's own reachability — not whether the worker's registered connection is actually polling — so it reports ok throughout an outage exactly like this one).

Suggested fix

Either (or both):

  1. Retry the initial connection with backoff inside onModuleInit()/initializeMultipleWorkers() instead of giving up after one attempt.
  2. Provide (or default to) a mode where a failed initial connection is treated as fatal (throws instead of silently returning), so the host process supervisor's own crash-restart behavior provides the retry loop for free.

Happy to provide the exact getTemporalModule() call site in Postiz's own code if useful — it never sets allowConnectionFailure, so it inherits this library's silently-tolerant default for worker registration specifically (not for the simpler client-only connections elsewhere in the app, which don't go through this same code path and were unaffected during this incident).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions