Skip to content

fix(liveness): gated peers start dead until confirmed reachable - #6

Merged
mxhob1 merged 1 commit into
noden/mainfrom
fix/liveness-start-dead
Jun 6, 2026
Merged

mxhob1 merged 1 commit into
noden/mainfrom
fix/liveness-start-dead

Conversation

@mxhob1

@mxhob1 mxhob1 commented Jun 6, 2026

Copy link
Copy Markdown
Member

Problem

When the liveness controller starts, gated peers (passive/active) defaulted to LIVE. A peer that is actually DOWN therefore kept its downstream spec.routes installed during the initial window (~30–45s) until down-detection finally fired. That blackholes the peer's subnet instead of letting cluster→tunnel traffic fall back to a broader live peer's covering route.

Observed in production (cluster default = active): a down site peer kept 10.254.2.0/24 installed for ~32s after agent start, blackholing that subnet instead of falling back to another peer's covering /16.

Fix

Gated peers now start DEAD and only become live once the liveness source positively confirms reachability. Two start-live sources were closed:

  • passive — observe() seeded lastAdvance off the first device read whenever ReceiveBytes > 0. wgctrl counters persist on the live wg0 across an agent pod restart, so a stale nonzero counter faked progress ⇒ live. The first observe now records only an RX baseline (seenRX); only a subsequent RX increase counts as relative progress. The absolute LastHandshakeTime branch is unchanged and still fires on tick 1, so a genuinely-connected peer is live on the first check.
  • active — IsLive returned failed < failures, and failed starts at 0, so an active peer was live before any probe/progress confirmed it. Active peers now track a confirmed flag, set the first time the peer is passive-live (fresh handshake / inbound progress) or answers a probe; IsLive requires confirmed && failed < failures.

disabled-mode peers are unchanged — the controller's IsLive short-circuits to true for disabled, so fallback gateways install their routes immediately with no initial gap.

Net effect

  • A down gated peer installs no routes at startup → traffic falls back to a broader live peer.
  • A live gated peer (recent handshake — absolute, evaluable on tick 1) gets its routes back within ~1 check interval (≈1s) of startup.

Tests

Added start-dead coverage for both passive and active: stale-counter/no-handshake stays dead and installs no routes; fresh-handshake is live on the first check; disabled peer is always live from the start. Updated the existing tests that relied on the old single-observe-equals-live behavior to seed a fresh handshake.

go build ./..., go vet ./..., and go test for the changed package pass; the golangci-lint-full pre-commit hook passed. (The internal/controller envtest suite is unaffected by this change and times out locally without KUBEBUILDER_ASSETS — environmental, not a regression.)

🤖 Generated with Claude Code

Gated (passive/active) peers were classified LIVE on controller start, so a
peer that is actually DOWN kept its downstream spec.routes installed during
the initial window (~30-45s) until down-detection finally fired. That
blackholed the peer's subnet instead of letting it fall back to a broader
live peer's covering route. Observed in prod (cluster default = active): a
down site peer held 10.254.2.0/24 for ~32s after agent start.

Two start-live sources fixed:

- passive: observe() seeded lastAdvance off the FIRST device read whenever
  ReceiveBytes > 0. wgctrl counters persist on the live wg0 across an agent
  pod restart, so a stale nonzero counter faked progress => live. Now the
  first observe records only an RX baseline (seenRX); only a SUBSEQUENT RX
  increase counts as relative progress. The absolute LastHandshakeTime branch
  is unchanged and still fires on tick 1, so a genuinely-connected peer is
  live on the first check.

- active: IsLive returned failed < failures, and failed starts at 0, so an
  active peer was live before any probe/progress confirmed it. Now active
  peers track a `confirmed` flag set the first time the peer is passive-live
  (fresh handshake / inbound progress) or answers a probe; IsLive requires
  confirmed && failed < failures.

disabled-mode peers are unchanged: the controller's IsLive short-circuits to
true for disabled, so fallback gateways install their routes immediately with
no initial gap.

Net effect: down gated peers install no routes at startup (traffic falls back
to a broader live peer); live gated peers get their routes within ~1 check
interval.

Tests: added start-dead coverage (stale-counter/no-handshake stays dead and
installs no routes; fresh-handshake live on first check; disabled always live)
for both passive and active; updated existing tests that relied on the old
single-observe-equals-live behavior to seed a fresh handshake.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@mxhob1
mxhob1 merged commit b805382 into noden/main Jun 6, 2026
5 checks passed
@mxhob1
mxhob1 deleted the fix/liveness-start-dead branch June 6, 2026 15:46
mxhob1 added a commit that referenced this pull request Jun 6, 2026
Routes declared on a WireguardPeer (spec.routes/routesV6) are installed
unconditionally today. This adds an agent-side liveness gate so a peer's
routes are withdrawn from both AllowedIPs and the kernel route table when the
peer goes unreachable, and re-attached when it recovers — letting cluster→
tunnel traffic fail over to a broader, still-live peer (longest-prefix).

Modes (cascade: per-peer spec.routeLiveness > per-instance spec.routeLiveness
> cluster default WG_ROUTE_LIVENESS env):
  - disabled — routes always installed (current static behaviour; default)
  - passive  — withdraw when inbound goes silent (ReceiveBytes/handshake stall
               over a per-peer N*keepalive window)
  - active   — passive + a /32 UDP handshake probe; down after N unanswered,
               revived on progress

Gated peers start dead until confirmed reachable, so a broader live peer wins
the longest-prefix match during the initial bring-up window rather than
blackholing through a peer that hasn't completed a handshake yet.

Design notes:
  - The /32 peer address is always kept in AllowedIPs; only spec.routes are
    gated, so the control channel to the peer never drops.
  - New env knobs are delivered via the agent deployment template (operator env
    passthrough), not the CRD: WG_ROUTE_LIVENESS, WG_ROUTE_FAILURE_COUNT,
    WG_ROUTE_CHECK_INTERVAL, WG_ROUTE_PROBE_INTERVAL. Unset ⇒ disabled ⇒
    byte-identical to current behaviour.
  - The controller reconciles these WG_ROUTE_* env vars onto existing agent
    Deployments (by name), so a Deployment created before the setting was
    applied (or before it changed) picks it up instead of staying inert; no
    churn once they match.
  - spec.routeLiveness is a plain optional string (not apiserver-enum) so a
    renderer can always emit ""=inherit; the agent validates (unknown ⇒
    inherit, fail-safe).
  - Adds a wireguard_peer_routes_active gauge + transition logging.

Squashed from internal PRs #4 (core), #5 (per-peer/instance cascade), #6
(start-dead) and #7 (reconcile env on existing Deployments). Depends on the
peer Routes/RoutesV6 feature.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant