Repository navigation
feat: liveness-gated route failover (disabled/passive/active) - #4
Merged
Merged
Conversation
…sive/active), agent-side extension of Phase-G route path
…viour-preservation invariants (Disabled peers + disabled mode unchanged vs noden/main)
…active down detection; add active-mode probe-interval (p) knob; consolidate Configuration section
… + active probe interval into one WG_ROUTE_INTERVAL (5s)
…lit from new WG_ROUTE_PROBE_INTERVAL 5s (rate-limited active probe); note both are net-new + non-CRD delivery via agent deployment template
…(read/check cadence; distinct from WG_ROUTE_PROBE_INTERVAL)
…MEOUT; active down ~45s)
…r (12 TDD tasks) + spec precondition (b) verified against KRO RGD
…, no behavior change)
…l, per-peer N*keepalive window)
…ered, progress revives
…ment-template passthrough)
…nused tickOnce return
mxhob1
added a commit
that referenced
this pull request
Jun 6, 2026
Routes declared on a WireguardPeer (spec.routes/routesV6) are installed
unconditionally today. This adds an agent-side liveness gate so a peer's
routes are withdrawn from both AllowedIPs and the kernel route table when the
peer goes unreachable, and re-attached when it recovers — letting cluster→
tunnel traffic fail over to a broader, still-live peer (longest-prefix).
Modes (cascade: per-peer spec.routeLiveness > per-instance spec.routeLiveness
> cluster default WG_ROUTE_LIVENESS env):
- disabled — routes always installed (current static behaviour; default)
- passive — withdraw when inbound goes silent (ReceiveBytes/handshake stall
over a per-peer N*keepalive window)
- active — passive + a /32 UDP handshake probe; down after N unanswered,
revived on progress
Gated peers start dead until confirmed reachable, so a broader live peer wins
the longest-prefix match during the initial bring-up window rather than
blackholing through a peer that hasn't completed a handshake yet.
Design notes:
- The /32 peer address is always kept in AllowedIPs; only spec.routes are
gated, so the control channel to the peer never drops.
- New env knobs are delivered via the agent deployment template (operator env
passthrough), not the CRD: WG_ROUTE_LIVENESS, WG_ROUTE_FAILURE_COUNT,
WG_ROUTE_CHECK_INTERVAL, WG_ROUTE_PROBE_INTERVAL. Unset ⇒ disabled ⇒
byte-identical to current behaviour.
- The controller reconciles these WG_ROUTE_* env vars onto existing agent
Deployments (by name), so a Deployment created before the setting was
applied (or before it changed) picks it up instead of staying inert; no
churn once they match.
- spec.routeLiveness is a plain optional string (not apiserver-enum) so a
renderer can always emit ""=inherit; the agent validates (unknown ⇒
inherit, fail-safe).
- Adds a wireguard_peer_routes_active gauge + transition logging.
Squashed from internal PRs #4 (core), #5 (per-peer/instance cascade), #6
(start-dead) and #7 (reconcile env on existing Deployments). Depends on the
peer Routes/RoutesV6 feature.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements
docs/superpowers/specs/2026-06-05-liveness-gated-route-failover-design.md(plan:docs/superpowers/plans/2026-06-06-liveness-gated-route-failover.md).A peer's downstream
spec.routesinstall inwg0only while the peer is reachable, so traffic fails over to a broader live peer (longest-prefix) when it goes silent and re-attaches when it recovers. Pure agent-side extension of the existing Phase-G route path — no controller/CRD/API changes.Mechanism
LivenessSourcepredicate threaded into the existing route filters (peerAllowedIPs,desiredKernelRoutes): a not-live peer is treated like aDisabledpeer for routes only — its/32stays so it can handshake back. All add/remove reuses the existingwg syncconfdiff +RouteReplace/RouteDelprune (no new install code).wgctrlcounters everyWG_ROUTE_CHECK_INTERVALand calls the existingwg.Synconly on a transition.Modes (
WG_ROUTE_LIVENESS, defaultdisabled)IsLive≡true, watcher not started → byte-identical to current behavior.ReceiveBytes/handshake stalls forN × peer.PersistentKeepalive(per-peer;180sREJECT_AFTER_TIME fallback for keepalive-less peers). Zero added traffic./32UDP handshake-probe everyWG_ROUTE_PROBE_INTERVAL; down afterNunanswered probes; any inbound progress revives. Closes the keepalive-less recovery deadlock.Config (all new;
os.Getenv, surfaced on the agent container via the deployment template)WG_ROUTE_LIVENESS=disabled|passive|active,WG_ROUTE_FAILURE_COUNT=3(N),WG_ROUTE_CHECK_INTERVAL=1s,WG_ROUTE_PROBE_INTERVAL=15s. Unset ⇒ disabled ⇒ no-op.Tests
TDD, table-driven (the 4 edge scenarios incl. never-active / went-silent / keepalive-less).
go build,go vet,golangci-lint-full, and all non-envtest unit packages green.internal/controller(envtest) +internal/it(e2e) run in CI.Observability
wireguard_peer_routes_active{peer}gauge + one log line per transition.🤖 Generated with Claude Code