The gap
Nothing in the company watches whether a bound connector is actually receiving.
Established 2026-08-15: the WhatsApp connector — connection conn_01KTW15C6ARW02CT4AS6CK0CF5, unarchived, active allow-list, bound to session sess_01KTVY2HE8B7J5EF20ZP1KEFY2 — has had no session activity since 2026-08-06. Nine days. Its container (aios-whatsapp) is crashlooping roughly once a minute.
The only thing that noticed was an unrelated audit, whose alarm was titled "audit cron down" — naming the detector rather than the finding — so it accumulated 29 comments pointing at the healthy component while a user-facing channel stayed dark.
Why the existing monitoring could not catch it
The company watches disk, model-pool utilization, deploy lag, PR gates, hold staleness, issue limbo, lane liveness, sweep freshness, migration chains. All of those watch things that PRODUCE a signal when unhealthy.
A silent connector produces nothing. And "no inbound messages" is exactly what a healthy, quiet channel looks like. Absence of inbound is not evidence of health — and today there is no signal that tells the two apart.
This is the same shape as the two lane reconcilers that sat dead for 15–18 days (eumemic-company#251): a healthy-looking system with nothing asserting it is alive. That one was fixed with a liveness detector; connectors have no equivalent.
Proposal
A connector-liveness check over every unarchived connection:
Thresholds should be per-channel — nine days of silence is alarming on an active support channel and meaningless on a rarely-used one. A single global threshold would produce exactly the noise that makes an alarm ignorable (see eumemic-ops#438).
Acceptance
Provenance: filed from the eumemic-ops#276 investigation, 2026-08-15.
The gap
Nothing in the company watches whether a bound connector is actually receiving.
Established 2026-08-15: the WhatsApp connector — connection
conn_01KTW15C6ARW02CT4AS6CK0CF5, unarchived, active allow-list, bound to sessionsess_01KTVY2HE8B7J5EF20ZP1KEFY2— has had no session activity since 2026-08-06. Nine days. Its container (aios-whatsapp) is crashlooping roughly once a minute.The only thing that noticed was an unrelated audit, whose alarm was titled "audit cron down" — naming the detector rather than the finding — so it accumulated 29 comments pointing at the healthy component while a user-facing channel stayed dark.
Why the existing monitoring could not catch it
The company watches disk, model-pool utilization, deploy lag, PR gates, hold staleness, issue limbo, lane liveness, sweep freshness, migration chains. All of those watch things that PRODUCE a signal when unhealthy.
This is the same shape as the two lane reconcilers that sat dead for 15–18 days (eumemic-company#251): a healthy-looking system with nothing asserting it is alive. That one was fixed with a liveness detector; connectors have no equivalent.
Proposal
A connector-liveness check over every unarchived connection:
Thresholds should be per-channel — nine days of silence is alarming on an active support channel and meaningless on a rarely-used one. A single global threshold would produce exactly the noise that makes an alarm ignorable (see eumemic-ops#438).
Acceptance
health_check_enabledis currently false onaios-whatsapp— worth asking whether every connector container should define what "healthy" means, since a restart loop with no success condition cannot converge.Provenance: filed from the eumemic-ops#276 investigation, 2026-08-15.