Skip to content

Nothing watches whether a bound connector is receiving — WhatsApp was dark 9 days and only an unrelated audit noticed #2153

Description

@eumemic

The gap

Nothing in the company watches whether a bound connector is actually receiving.

Established 2026-08-15: the WhatsApp connector — connection conn_01KTW15C6ARW02CT4AS6CK0CF5, unarchived, active allow-list, bound to session sess_01KTVY2HE8B7J5EF20ZP1KEFY2 — has had no session activity since 2026-08-06. Nine days. Its container (aios-whatsapp) is crashlooping roughly once a minute.

The only thing that noticed was an unrelated audit, whose alarm was titled "audit cron down" — naming the detector rather than the finding — so it accumulated 29 comments pointing at the healthy component while a user-facing channel stayed dark.

Why the existing monitoring could not catch it

The company watches disk, model-pool utilization, deploy lag, PR gates, hold staleness, issue limbo, lane liveness, sweep freshness, migration chains. All of those watch things that PRODUCE a signal when unhealthy.

A silent connector produces nothing. And "no inbound messages" is exactly what a healthy, quiet channel looks like. Absence of inbound is not evidence of health — and today there is no signal that tells the two apart.

This is the same shape as the two lane reconcilers that sat dead for 15–18 days (eumemic-company#251): a healthy-looking system with nothing asserting it is alive. That one was fixed with a liveness detector; connectors have no equivalent.

Proposal

A connector-liveness check over every unarchived connection:

  • For each bound connection, compare its container/process state against its bound session's last activity.
  • Alarm on the CONJUNCTION, which is the part that makes it low-noise: the transport is unhealthy AND the session has been silent beyond a per-channel threshold. Either alone is normal — a quiet channel is fine, and a brief restart is fine. Together they mean the channel is dark and nobody knows.
  • Report the finding, not the detector: "whatsapp: container restarting, no session activity in 9d".

Thresholds should be per-channel — nine days of silence is alarming on an active support channel and meaningless on a rarely-used one. A single global threshold would produce exactly the noise that makes an alarm ignorable (see eumemic-ops#438).

Acceptance

  • A connector whose container is unhealthy and whose session has been silent past its threshold produces an alarm naming both facts
  • A quiet-but-healthy channel produces nothing — the true-negative, and the thing that decides whether this detector gets read or ignored
  • Mutation test: stop a connector container in a test environment → the alarm fires. An untested detector is how the lane reconcilers went 18 days unnoticed.
  • health_check_enabled is currently false on aios-whatsapp — worth asking whether every connector container should define what "healthy" means, since a restart loop with no success condition cannot converge.

Provenance: filed from the eumemic-ops#276 investigation, 2026-08-15.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    approvedGreenlit to build. Dispatch gate = shovel-ready + approved.enhancementNew feature or requestpriority:highshovel-readyDesign settled, scope clear; ready to implement without further design discussion

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions