Use a startup taint plus a readiness controller. The controller
should remove the taint only when every DaemonSet explicitly designated
as required has a Ready pod on that node.
This keeps Kubernetes’ native NodeReady condition separate from
workload initialization and prevents ordinary workloads from scheduling
too early. Avoid relying on a fixed sleep, which can fail when images
take longer to pull or a DaemonSet is unhealthy.
The implementation consists of:
NodePool startup taint.Merge this into your existing NodePool; retain your existing
nodeClassRef, requirements, limits, and disruption settings.
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: default
spec:
template:
spec:
startupTaints:
- key: readiness.example.com/required-daemonsets
value: pending
effect: NoSchedule
Karpenter uses startup taints for temporary initialization requirements
and expects another component to remove them.
Important: Apply this only to the relevant NodePools. Updating a
NodePool does not necessarily add the startup taint to nodes that
already exist. Verify how your existing nodes are managed before rolling
out the change.
Documentation: Karpenter
NodePools
Use the project’s official release manifests. The version below is an
example from the installation guide; check the releases
page
and current installation instructions before deploying to production.
VERSION=v0.5.0
kubectl apply -f \
"https://github.com/kubernetes-sigs/node-readiness-controller/releases/download/${VERSION}/crds.yaml"
kubectl wait \
--for=condition=Established \
--timeout=60s \
crd/nodereadinessrules.readiness.node.x-k8s.io
kubectl apply -f \
"https://github.com/kubernetes-sigs/node-readiness-controller/releases/download/${VERSION}/install.yaml"
Project: Kubernetes SIGs Node Readiness
Controller
The rule requires a custom node condition named
readiness.example.com/RequiredDaemonSetsReady to be True. The
reporter described in the next section is responsible for setting this
condition.
apiVersion: readiness.node.x-k8s.io/v1alpha1
kind: NodeReadinessRule
metadata:
name: required-daemonsets-ready
spec:
nodeSelector:
matchLabels:
readiness.example.com/enabled: "true"
conditions:
- type: readiness.example.com/RequiredDaemonSetsReady
requiredStatus: "True"
taint:
key: readiness.example.com/required-daemonsets
value: pending
effect: NoSchedule
enforcementMode: continuous
The node selector means that only nodes labelled
readiness.example.com/enabled=true are governed by this rule.
Continuous enforcement is intended to restore the taint if the required
readiness condition later becomes false. A NoSchedule taint blocks new
workloads but does not automatically evict workloads already running on
the node.
Check the current API schema and examples in the Node Readiness
Controller
documentation
before applying this manifest.
Use an explicit allowlist of DaemonSets that must be ready on every
applicable node.
apiVersion: v1
kind: ConfigMap
metadata:
name: daemonset-readiness-config
namespace: kube-system
data:
required-daemonsets.yaml: |
required:
- namespace: kube-system
name: kube-proxy
- namespace: kube-system
name: aws-node
- namespace: kube-system
name: ebs-csi-node
- namespace: monitoring
name: node-exporter
Replace these examples with the DaemonSets actually required in your
cluster. For example, if you use Cilium instead of the AWS VPC CNI,
replace aws-node with your Cilium DaemonSet.
The Node Readiness Controller does not automatically inspect arbitrary
DaemonSet pod readiness. A reporter must check the required pods and
update the custom node condition.
For each node, the reporter should:
Ready=True.readiness.example.com/RequiredDaemonSetsReady=True only whenFalse if a required pod is missing, unready,Treat a missing or unknown readiness signal as not ready.
The following is the RBAC portion for a reporter that reads pods and
DaemonSets and updates node status. It is not a complete deployable
reporter.
apiVersion: v1
kind: ServiceAccount
metadata:
name: daemonset-readiness-reporter
namespace: kube-system
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: daemonset-readiness-reporter
rules:
- apiGroups: [""]
resources: ["pods", "nodes"]
verbs: ["get", "list", "watch"]
- apiGroups: ["apps"]
resources: ["daemonsets"]
verbs: ["get", "list", "watch"]
- apiGroups: [""]
resources: ["nodes/status"]
verbs: ["patch", "update"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: daemonset-readiness-reporter
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: daemonset-readiness-reporter
subjects:
- kind: ServiceAccount
name: daemonset-readiness-reporter
namespace: kube-system
You must still deploy reporter code as a DaemonSet, provide the node
name through the downward API, and implement the condition-update loop.
The allowlist is illustrative; do not use it unchanged if those
DaemonSets are not required in your cluster.
Every required DaemonSet must tolerate the custom NoSchedule taint;
otherwise, its pod may never schedule and the node may remain blocked
indefinitely.
Add this toleration to each required DaemonSet’s pod template:
spec:
template:
spec:
tolerations:
- key: readiness.example.com/required-daemonsets
operator: Equal
value: pending
effect: NoSchedule
The readiness reporter must tolerate the same taint. The Node Readiness
Controller must remain schedulable on an existing healthy node so that
it can remove taints from newly provisioned nodes.
After deploying the configuration and reporter, inspect the state:
kubectl get nodepools
kubectl get nodes
kubectl get nodereadinessrules
kubectl get nodes \
-o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints
Test on a non-production NodePool:
True.NoSchedule taint blocks new scheduling; it does not evict
-
https://github.com/kubernetes-sigs/node-readiness-controller