4370 links
  • Arnaud's links
  • Home
  • Login
  • RSS Feed
  • ATOM Feed
  • Tag cloud
  • Picture wall
  • Daily
Links per page: 20 50 100
  • thumbnail
    GitHub - kubernetes-sigs/node-readiness-controller: Declarative node readiness for Kubernetes - gate scheduling with taints until node infrastructure (CNI, GPU, storage, custom checks) is actually ready. · GitHub

    Karpenter: Make Nodes Ready Only After Required DaemonSets Are Ready

    Recommended approach

    Use a startup taint plus a readiness controller. The controller
    should remove the taint only when every DaemonSet explicitly designated
    as required has a Ready pod on that node.

    This keeps Kubernetes’ native NodeReady condition separate from
    workload initialization and prevents ordinary workloads from scheduling
    too early. Avoid relying on a fixed sleep, which can fail when images
    take longer to pull or a DaemonSet is unhealthy.

    The implementation consists of:

    1. A Karpenter NodePool startup taint.
    2. An allowlist of required DaemonSets.
    3. A readiness reporter that checks the pods on each node and publishes
      a custom node condition.
    4. A Node Readiness Controller rule that removes the taint when the
      condition is satisfied.

    1. Add a startup taint to the Karpenter NodePool

    Merge this into your existing NodePool; retain your existing
    nodeClassRef, requirements, limits, and disruption settings.

    apiVersion: karpenter.sh/v1
    kind: NodePool
    metadata:
      name: default
    spec:
      template:
        spec:
          startupTaints:
            - key: readiness.example.com/required-daemonsets
              value: pending
              effect: NoSchedule

    Karpenter uses startup taints for temporary initialization requirements
    and expects another component to remove them.

    Important: Apply this only to the relevant NodePools. Updating a
    NodePool does not necessarily add the startup taint to nodes that
    already exist. Verify how your existing nodes are managed before rolling
    out the change.

    Documentation: Karpenter
    NodePools

    2. Install the Node Readiness Controller

    Use the project’s official release manifests. The version below is an
    example from the installation guide; check the releases
    page
    and current installation instructions before deploying to production.

    VERSION=v0.5.0
    
    kubectl apply -f \
      "https://github.com/kubernetes-sigs/node-readiness-controller/releases/download/${VERSION}/crds.yaml"
    
    kubectl wait \
      --for=condition=Established \
      --timeout=60s \
      crd/nodereadinessrules.readiness.node.x-k8s.io
    
    kubectl apply -f \
      "https://github.com/kubernetes-sigs/node-readiness-controller/releases/download/${VERSION}/install.yaml"

    Project: Kubernetes SIGs Node Readiness
    Controller

    3. Define the readiness rule

    The rule requires a custom node condition named
    readiness.example.com/RequiredDaemonSetsReady to be True. The
    reporter described in the next section is responsible for setting this
    condition.

    apiVersion: readiness.node.x-k8s.io/v1alpha1
    kind: NodeReadinessRule
    metadata:
      name: required-daemonsets-ready
    spec:
      nodeSelector:
        matchLabels:
          readiness.example.com/enabled: "true"
    
      conditions:
        - type: readiness.example.com/RequiredDaemonSetsReady
          requiredStatus: "True"
    
      taint:
        key: readiness.example.com/required-daemonsets
        value: pending
        effect: NoSchedule
    
      enforcementMode: continuous

    The node selector means that only nodes labelled
    readiness.example.com/enabled=true are governed by this rule.

    Continuous enforcement is intended to restore the taint if the required
    readiness condition later becomes false. A NoSchedule taint blocks new
    workloads but does not automatically evict workloads already running on
    the node.

    Check the current API schema and examples in the Node Readiness
    Controller
    documentation
    before applying this manifest.

    4. Define the required DaemonSets

    Use an explicit allowlist of DaemonSets that must be ready on every
    applicable node.

    apiVersion: v1
    kind: ConfigMap
    metadata:
      name: daemonset-readiness-config
      namespace: kube-system
    data:
      required-daemonsets.yaml: |
        required:
          - namespace: kube-system
            name: kube-proxy
          - namespace: kube-system
            name: aws-node
          - namespace: kube-system
            name: ebs-csi-node
          - namespace: monitoring
            name: node-exporter

    Replace these examples with the DaemonSets actually required in your
    cluster. For example, if you use Cilium instead of the AWS VPC CNI,
    replace aws-node with your Cilium DaemonSet.

    5. Implement the readiness reporter

    The Node Readiness Controller does not automatically inspect arbitrary
    DaemonSet pod readiness. A reporter must check the required pods and
    update the custom node condition.

    For each node, the reporter should:

    1. Read the node name from the downward API.
    2. Check each required DaemonSet’s pod on that node.
    3. Confirm that each pod is owned by the expected DaemonSet and has
      Ready=True.
    4. Set readiness.example.com/RequiredDaemonSetsReady=True only when
      every required DaemonSet has a ready pod on that node.
    5. Set the condition to False if a required pod is missing, unready,
      or terminating.

    Treat a missing or unknown readiness signal as not ready.

    Reporter RBAC

    The following is the RBAC portion for a reporter that reads pods and
    DaemonSets and updates node status. It is not a complete deployable
    reporter.

    apiVersion: v1
    kind: ServiceAccount
    metadata:
      name: daemonset-readiness-reporter
      namespace: kube-system
    ---
    apiVersion: rbac.authorization.k8s.io/v1
    kind: ClusterRole
    metadata:
      name: daemonset-readiness-reporter
    rules:
      - apiGroups: [""]
        resources: ["pods", "nodes"]
        verbs: ["get", "list", "watch"]
      - apiGroups: ["apps"]
        resources: ["daemonsets"]
        verbs: ["get", "list", "watch"]
      - apiGroups: [""]
        resources: ["nodes/status"]
        verbs: ["patch", "update"]
    ---
    apiVersion: rbac.authorization.k8s.io/v1
    kind: ClusterRoleBinding
    metadata:
      name: daemonset-readiness-reporter
    roleRef:
      apiGroup: rbac.authorization.k8s.io
      kind: ClusterRole
      name: daemonset-readiness-reporter
    subjects:
      - kind: ServiceAccount
        name: daemonset-readiness-reporter
        namespace: kube-system

    You must still deploy reporter code as a DaemonSet, provide the node
    name through the downward API, and implement the condition-update loop.
    The allowlist is illustrative; do not use it unchanged if those
    DaemonSets are not required in your cluster.

    6. Make required DaemonSets tolerate the startup taint

    Every required DaemonSet must tolerate the custom NoSchedule taint;
    otherwise, its pod may never schedule and the node may remain blocked
    indefinitely.

    Add this toleration to each required DaemonSet’s pod template:

    spec:
      template:
        spec:
          tolerations:
            - key: readiness.example.com/required-daemonsets
              operator: Equal
              value: pending
              effect: NoSchedule

    The readiness reporter must tolerate the same taint. The Node Readiness
    Controller must remain schedulable on an existing healthy node so that
    it can remove taints from newly provisioned nodes.

    7. Validate the rollout

    After deploying the configuration and reporter, inspect the state:

    kubectl get nodepools
    kubectl get nodes
    kubectl get nodereadinessrules
    kubectl get nodes \
      -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints

    Test on a non-production NodePool:

    • Provision a new node.
    • Confirm the startup taint prevents ordinary workloads from
      scheduling.
    • Confirm all required DaemonSet pods become ready.
    • Confirm the reporter sets the custom node condition to True.
    • Confirm the taint is removed.
    • Stop a required DaemonSet pod and verify that continuous enforcement
      restores the taint.

    Important limitations

    • A NoSchedule taint blocks new scheduling; it does not evict
      workloads already running on the node.
    • Evicting existing workloads when a required DaemonSet becomes
      unready needs a separate policy and should be designed carefully to
      avoid unnecessary disruption.
    • Validate the Node Readiness Controller’s current CRD schema, release
      version, and rule semantics against its official documentation
      before applying the sample rule.
    • The reporter is a required implementation component; the YAML above
      provides its configuration and RBAC starting point, not the
      reporter’s executable code.

    References

    • Kubernetes SIGs Node Readiness
      Controller
    • Node Readiness Controller: getting
      started
    • Node Readiness Controller: CNI readiness
      example
    • Karpenter NodePools
      documentation
    October 11, 2026 at 7:45:51 AM GMT+2 * - permalink - archive.org - https://github.com/kubernetes-sigs/node-readiness-controller
    k8s node ready
Links per page: 20 50 100
Shaarli - The personal, minimalist, super fast, database-free, bookmarking service by the Shaarli community - Help/documentation