Repository navigation
fix(liveness): gated peers start dead until confirmed reachable - #6
Merged
Merged
Conversation
Gated (passive/active) peers were classified LIVE on controller start, so a peer that is actually DOWN kept its downstream spec.routes installed during the initial window (~30-45s) until down-detection finally fired. That blackholed the peer's subnet instead of letting it fall back to a broader live peer's covering route. Observed in prod (cluster default = active): a down site peer held 10.254.2.0/24 for ~32s after agent start. Two start-live sources fixed: - passive: observe() seeded lastAdvance off the FIRST device read whenever ReceiveBytes > 0. wgctrl counters persist on the live wg0 across an agent pod restart, so a stale nonzero counter faked progress => live. Now the first observe records only an RX baseline (seenRX); only a SUBSEQUENT RX increase counts as relative progress. The absolute LastHandshakeTime branch is unchanged and still fires on tick 1, so a genuinely-connected peer is live on the first check. - active: IsLive returned failed < failures, and failed starts at 0, so an active peer was live before any probe/progress confirmed it. Now active peers track a `confirmed` flag set the first time the peer is passive-live (fresh handshake / inbound progress) or answers a probe; IsLive requires confirmed && failed < failures. disabled-mode peers are unchanged: the controller's IsLive short-circuits to true for disabled, so fallback gateways install their routes immediately with no initial gap. Net effect: down gated peers install no routes at startup (traffic falls back to a broader live peer); live gated peers get their routes within ~1 check interval. Tests: added start-dead coverage (stale-counter/no-handshake stays dead and installs no routes; fresh-handshake live on first check; disabled always live) for both passive and active; updated existing tests that relied on the old single-observe-equals-live behavior to seed a fresh handshake. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mxhob1
added a commit
that referenced
this pull request
Jun 6, 2026
Routes declared on a WireguardPeer (spec.routes/routesV6) are installed
unconditionally today. This adds an agent-side liveness gate so a peer's
routes are withdrawn from both AllowedIPs and the kernel route table when the
peer goes unreachable, and re-attached when it recovers — letting cluster→
tunnel traffic fail over to a broader, still-live peer (longest-prefix).
Modes (cascade: per-peer spec.routeLiveness > per-instance spec.routeLiveness
> cluster default WG_ROUTE_LIVENESS env):
- disabled — routes always installed (current static behaviour; default)
- passive — withdraw when inbound goes silent (ReceiveBytes/handshake stall
over a per-peer N*keepalive window)
- active — passive + a /32 UDP handshake probe; down after N unanswered,
revived on progress
Gated peers start dead until confirmed reachable, so a broader live peer wins
the longest-prefix match during the initial bring-up window rather than
blackholing through a peer that hasn't completed a handshake yet.
Design notes:
- The /32 peer address is always kept in AllowedIPs; only spec.routes are
gated, so the control channel to the peer never drops.
- New env knobs are delivered via the agent deployment template (operator env
passthrough), not the CRD: WG_ROUTE_LIVENESS, WG_ROUTE_FAILURE_COUNT,
WG_ROUTE_CHECK_INTERVAL, WG_ROUTE_PROBE_INTERVAL. Unset ⇒ disabled ⇒
byte-identical to current behaviour.
- The controller reconciles these WG_ROUTE_* env vars onto existing agent
Deployments (by name), so a Deployment created before the setting was
applied (or before it changed) picks it up instead of staying inert; no
churn once they match.
- spec.routeLiveness is a plain optional string (not apiserver-enum) so a
renderer can always emit ""=inherit; the agent validates (unknown ⇒
inherit, fail-safe).
- Adds a wireguard_peer_routes_active gauge + transition logging.
Squashed from internal PRs #4 (core), #5 (per-peer/instance cascade), #6
(start-dead) and #7 (reconcile env on existing Deployments). Depends on the
peer Routes/RoutesV6 feature.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
When the liveness controller starts, gated peers (passive/active) defaulted to LIVE. A peer that is actually DOWN therefore kept its downstream
spec.routesinstalled during the initial window (~30–45s) until down-detection finally fired. That blackholes the peer's subnet instead of letting cluster→tunnel traffic fall back to a broader live peer's covering route.Observed in production (cluster default =
active): a down site peer kept10.254.2.0/24installed for ~32s after agent start, blackholing that subnet instead of falling back to another peer's covering/16.Fix
Gated peers now start DEAD and only become live once the liveness source positively confirms reachability. Two start-live sources were closed:
observe()seededlastAdvanceoff the first device read wheneverReceiveBytes > 0. wgctrl counters persist on the livewg0across an agent pod restart, so a stale nonzero counter faked progress ⇒ live. The first observe now records only an RX baseline (seenRX); only a subsequent RX increase counts as relative progress. The absoluteLastHandshakeTimebranch is unchanged and still fires on tick 1, so a genuinely-connected peer is live on the first check.IsLivereturnedfailed < failures, andfailedstarts at0, so an active peer was live before any probe/progress confirmed it. Active peers now track aconfirmedflag, set the first time the peer is passive-live (fresh handshake / inbound progress) or answers a probe;IsLiverequiresconfirmed && failed < failures.disabled-mode peers are unchanged — the controller'sIsLiveshort-circuits totruefor disabled, so fallback gateways install their routes immediately with no initial gap.Net effect
Tests
Added start-dead coverage for both passive and active: stale-counter/no-handshake stays dead and installs no routes; fresh-handshake is live on the first check;
disabledpeer is always live from the start. Updated the existing tests that relied on the old single-observe-equals-live behavior to seed a fresh handshake.go build ./...,go vet ./..., andgo testfor the changed package pass; thegolangci-lint-fullpre-commit hook passed. (Theinternal/controllerenvtest suite is unaffected by this change and times out locally withoutKUBEBUILDER_ASSETS— environmental, not a regression.)🤖 Generated with Claude Code