Skip to content

feat: liveness-gated route failover (disabled/passive/active) - #4

Merged
mxhob1 merged 20 commits into
noden/mainfrom
feat/liveness-gated-routes
Jun 6, 2026
Merged

mxhob1 merged 20 commits into
noden/mainfrom
feat/liveness-gated-routes

Conversation

@mxhob1

@mxhob1 mxhob1 commented Jun 6, 2026

Copy link
Copy Markdown
Member

Implements docs/superpowers/specs/2026-06-05-liveness-gated-route-failover-design.md (plan: docs/superpowers/plans/2026-06-06-liveness-gated-route-failover.md).

A peer's downstream spec.routes install in wg0 only while the peer is reachable, so traffic fails over to a broader live peer (longest-prefix) when it goes silent and re-attaches when it recovers. Pure agent-side extension of the existing Phase-G route path — no controller/CRD/API changes.

Mechanism

  • LivenessSource predicate threaded into the existing route filters (peerAllowedIPs, desiredKernelRoutes): a not-live peer is treated like a Disabled peer for routes only — its /32 stays so it can handshake back. All add/remove reuses the existing wg syncconf diff + RouteReplace/RouteDel prune (no new install code).
  • A watcher reads wgctrl counters every WG_ROUTE_CHECK_INTERVAL and calls the existing wg.Sync only on a transition.

Modes (WG_ROUTE_LIVENESS, default disabled)

  • disabled — IsLive≡true, watcher not started → byte-identical to current behavior.
  • passive — down when ReceiveBytes/handshake stalls for N × peer.PersistentKeepalive (per-peer; 180s REJECT_AFTER_TIME fallback for keepalive-less peers). Zero added traffic.
  • active — passive + a /32 UDP handshake-probe every WG_ROUTE_PROBE_INTERVAL; down after N unanswered probes; any inbound progress revives. Closes the keepalive-less recovery deadlock.

Config (all new; os.Getenv, surfaced on the agent container via the deployment template)

WG_ROUTE_LIVENESS=disabled|passive|active, WG_ROUTE_FAILURE_COUNT=3 (N), WG_ROUTE_CHECK_INTERVAL=1s, WG_ROUTE_PROBE_INTERVAL=15s. Unset ⇒ disabled ⇒ no-op.

Tests

TDD, table-driven (the 4 edge scenarios incl. never-active / went-silent / keepalive-less). go build, go vet, golangci-lint-full, and all non-envtest unit packages green. internal/controller (envtest) + internal/it (e2e) run in CI.

Observability

wireguard_peer_routes_active{peer} gauge + one log line per transition.

🤖 Generated with Claude Code

mxhob1 added 20 commits June 5, 2026 23:38
…sive/active), agent-side extension of Phase-G route path
…viour-preservation invariants (Disabled peers + disabled mode unchanged vs noden/main)
…active down detection; add active-mode probe-interval (p) knob; consolidate Configuration section
… + active probe interval into one WG_ROUTE_INTERVAL (5s)
…lit from new WG_ROUTE_PROBE_INTERVAL 5s (rate-limited active probe); note both are net-new + non-CRD delivery via agent deployment template
…(read/check cadence; distinct from WG_ROUTE_PROBE_INTERVAL)
…r (12 TDD tasks) + spec precondition (b) verified against KRO RGD
@mxhob1
mxhob1 merged commit baeffe9 into noden/main Jun 6, 2026
7 of 8 checks passed
mxhob1 added a commit that referenced this pull request Jun 6, 2026
Routes declared on a WireguardPeer (spec.routes/routesV6) are installed
unconditionally today. This adds an agent-side liveness gate so a peer's
routes are withdrawn from both AllowedIPs and the kernel route table when the
peer goes unreachable, and re-attached when it recovers — letting cluster→
tunnel traffic fail over to a broader, still-live peer (longest-prefix).

Modes (cascade: per-peer spec.routeLiveness > per-instance spec.routeLiveness
> cluster default WG_ROUTE_LIVENESS env):
  - disabled — routes always installed (current static behaviour; default)
  - passive  — withdraw when inbound goes silent (ReceiveBytes/handshake stall
               over a per-peer N*keepalive window)
  - active   — passive + a /32 UDP handshake probe; down after N unanswered,
               revived on progress

Gated peers start dead until confirmed reachable, so a broader live peer wins
the longest-prefix match during the initial bring-up window rather than
blackholing through a peer that hasn't completed a handshake yet.

Design notes:
  - The /32 peer address is always kept in AllowedIPs; only spec.routes are
    gated, so the control channel to the peer never drops.
  - New env knobs are delivered via the agent deployment template (operator env
    passthrough), not the CRD: WG_ROUTE_LIVENESS, WG_ROUTE_FAILURE_COUNT,
    WG_ROUTE_CHECK_INTERVAL, WG_ROUTE_PROBE_INTERVAL. Unset ⇒ disabled ⇒
    byte-identical to current behaviour.
  - The controller reconciles these WG_ROUTE_* env vars onto existing agent
    Deployments (by name), so a Deployment created before the setting was
    applied (or before it changed) picks it up instead of staying inert; no
    churn once they match.
  - spec.routeLiveness is a plain optional string (not apiserver-enum) so a
    renderer can always emit ""=inherit; the agent validates (unknown ⇒
    inherit, fail-safe).
  - Adds a wireguard_peer_routes_active gauge + transition logging.

Squashed from internal PRs #4 (core), #5 (per-peer/instance cascade), #6
(start-dead) and #7 (reconcile env on existing Deployments). Depends on the
peer Routes/RoutesV6 feature.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant