Skip to content

KopsControlPlane readiness flaps/stays false on high-churn Karpenter clusters #137

Description

@csmartins

Summary

On clusters with continuous Karpenter node churn, KopsControlPlane.status.ready frequently reports false even though the control plane and all nodes are healthy. The flag is derived from whether every machine has joined the cluster at the moment of reconcile — so any node that happens to be mid-registration (normal Karpenter behavior) flips the KCP to not-ready, and because the flag is only re-evaluated on the next reconcile (~15–21 min on a large fleet), it lingers or effectively never clears on the busiest clusters.

Environment

  • kops-operator: v0.23.2-alpha (vendored kops 1.36.1)
  • Workers: WORKER_COUNT=3
  • Node provisioning: Karpenter (fleet-wide), steady node churn
  • ~16 KopsControlPlanes reconciled by one operator instance

Observed behavior

  • KopsControlPlane.status.ready = false
  • All sub-conditions True: KopsControlPlaneSecretsReady, KopsControlPlaneStateReady,
    KopsTerraformGenerationReady, TerraformApplyReady
  • Terraform applied logged successfully; reconcile ends with an event
    machine "<id>" has not yet joined cluster
  • All NodeClaims Registered=True/Ready=True, no genuinely-stuck nodes; the node the
    reconcile flagged has since joined
  • The ready=false clears only on the next reconcile that happens to run while no node is
    mid-registration — on busy clusters that alignment rarely happens, so 1–2 clusters are
    perpetually false and the set rotates.

Root cause (suspected)

Readiness is computed as a strict "all machines joined" snapshot at reconcile time and
persisted to the top-level ready field, with no tolerance for nodes still within their
normal registration window and no re-evaluation between reconciles.

Impact

  • False-negative KCP readiness on the busiest clusters → noisy/misleading alerting.
  • Masks genuinely not-ready control planes (can't distinguish "a node is normally
    registering" from "the control plane is broken").
  • Prevents a clean "all-green" readout across a large fleet.

Suggested fix

  • Don't fail readiness for nodes still within an expected registration grace period (e.g.
    align with Karpenter's node registration TTL); only gate on nodes that have exceeded it.
  • And/or re-evaluate/clear the top-level ready flag more promptly than the full reconcile
    interval once machines have joined.
  • Distinguish "node normally registering" from "node genuinely stuck / control plane
    unhealthy".

References

  • Controller readiness/"has not yet joined" logic in controllers/controlplane/.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions