Use a startup taint plus a readiness controller. The controller
should remove the taint only when every DaemonSet explicitly designated
as required has a Ready pod on that node.
This keeps Kubernetes’ native NodeReady condition separate from
workload initialization and prevents ordinary workloads from scheduling
too early. Avoid relying on a fixed sleep, which can fail when images
take longer to pull or a DaemonSet is unhealthy.
The implementation consists of:
NodePool startup taint.Merge this into your existing NodePool; retain your existing
nodeClassRef, requirements, limits, and disruption settings.
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: default
spec:
template:
spec:
startupTaints:
- key: readiness.example.com/required-daemonsets
value: pending
effect: NoSchedule
Karpenter uses startup taints for temporary initialization requirements
and expects another component to remove them.
Important: Apply this only to the relevant NodePools. Updating a
NodePool does not necessarily add the startup taint to nodes that
already exist. Verify how your existing nodes are managed before rolling
out the change.
Documentation: Karpenter
NodePools
Use the project’s official release manifests. The version below is an
example from the installation guide; check the releases
page
and current installation instructions before deploying to production.
VERSION=v0.5.0
kubectl apply -f \
"https://github.com/kubernetes-sigs/node-readiness-controller/releases/download/${VERSION}/crds.yaml"
kubectl wait \
--for=condition=Established \
--timeout=60s \
crd/nodereadinessrules.readiness.node.x-k8s.io
kubectl apply -f \
"https://github.com/kubernetes-sigs/node-readiness-controller/releases/download/${VERSION}/install.yaml"
Project: Kubernetes SIGs Node Readiness
Controller
The rule requires a custom node condition named
readiness.example.com/RequiredDaemonSetsReady to be True. The
reporter described in the next section is responsible for setting this
condition.
apiVersion: readiness.node.x-k8s.io/v1alpha1
kind: NodeReadinessRule
metadata:
name: required-daemonsets-ready
spec:
nodeSelector:
matchLabels:
readiness.example.com/enabled: "true"
conditions:
- type: readiness.example.com/RequiredDaemonSetsReady
requiredStatus: "True"
taint:
key: readiness.example.com/required-daemonsets
value: pending
effect: NoSchedule
enforcementMode: continuous
The node selector means that only nodes labelled
readiness.example.com/enabled=true are governed by this rule.
Continuous enforcement is intended to restore the taint if the required
readiness condition later becomes false. A NoSchedule taint blocks new
workloads but does not automatically evict workloads already running on
the node.
Check the current API schema and examples in the Node Readiness
Controller
documentation
before applying this manifest.
Use an explicit allowlist of DaemonSets that must be ready on every
applicable node.
apiVersion: v1
kind: ConfigMap
metadata:
name: daemonset-readiness-config
namespace: kube-system
data:
required-daemonsets.yaml: |
required:
- namespace: kube-system
name: kube-proxy
- namespace: kube-system
name: aws-node
- namespace: kube-system
name: ebs-csi-node
- namespace: monitoring
name: node-exporter
Replace these examples with the DaemonSets actually required in your
cluster. For example, if you use Cilium instead of the AWS VPC CNI,
replace aws-node with your Cilium DaemonSet.
The Node Readiness Controller does not automatically inspect arbitrary
DaemonSet pod readiness. A reporter must check the required pods and
update the custom node condition.
For each node, the reporter should:
Ready=True.readiness.example.com/RequiredDaemonSetsReady=True only whenFalse if a required pod is missing, unready,Treat a missing or unknown readiness signal as not ready.
The following is the RBAC portion for a reporter that reads pods and
DaemonSets and updates node status. It is not a complete deployable
reporter.
apiVersion: v1
kind: ServiceAccount
metadata:
name: daemonset-readiness-reporter
namespace: kube-system
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: daemonset-readiness-reporter
rules:
- apiGroups: [""]
resources: ["pods", "nodes"]
verbs: ["get", "list", "watch"]
- apiGroups: ["apps"]
resources: ["daemonsets"]
verbs: ["get", "list", "watch"]
- apiGroups: [""]
resources: ["nodes/status"]
verbs: ["patch", "update"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: daemonset-readiness-reporter
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: daemonset-readiness-reporter
subjects:
- kind: ServiceAccount
name: daemonset-readiness-reporter
namespace: kube-system
You must still deploy reporter code as a DaemonSet, provide the node
name through the downward API, and implement the condition-update loop.
The allowlist is illustrative; do not use it unchanged if those
DaemonSets are not required in your cluster.
Every required DaemonSet must tolerate the custom NoSchedule taint;
otherwise, its pod may never schedule and the node may remain blocked
indefinitely.
Add this toleration to each required DaemonSet’s pod template:
spec:
template:
spec:
tolerations:
- key: readiness.example.com/required-daemonsets
operator: Equal
value: pending
effect: NoSchedule
The readiness reporter must tolerate the same taint. The Node Readiness
Controller must remain schedulable on an existing healthy node so that
it can remove taints from newly provisioned nodes.
After deploying the configuration and reporter, inspect the state:
kubectl get nodepools
kubectl get nodes
kubectl get nodereadinessrules
kubectl get nodes \
-o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints
Test on a non-production NodePool:
True.NoSchedule taint blocks new scheduling; it does not evict
-
https://github.com/kubernetes-sigs/node-readiness-controlleraws ecr get-account-setting --name REGISTRY_POLICY_SCOPE
aws ecr put-account-setting --name REGISTRY_POLICY_SCOPE --value V2
-
https://aws.amazon.com/about-aws/whats-new/2024/12/amazon-ecr-expands-registry-policy-ecr-actions/A long-awaited Karpenter feature is finally here!
The latest Karpenter release supports Capacity Buffer, eliminating the need for workarounds such as balloon pods to maintain spare node capacity.
A Capacity Buffer defines virtual placeholder pods that exist only in Karpenter’s scheduling simulation — they are never created as actual Kubernetes pods.
These virtual pods tell Karpenter to provision nodes with spare capacity. They participate in each scheduling cycle to maintain the buffer and are automatically refilled as real workloads consume the pre-provisioned capacity.
Benefits over previous workarounds:
• No balloon pods: Eliminates the need for low-priority placeholder deployments and complex PriorityClass configurations.
• More efficient scaling: Capacity can be defined using fixed counts or percentages, allowing spare capacity to scale with your workload instead of over-provisioning entire node pools.
• Automatic replenishment: The buffer is automatically maintained as workloads consume the available capacity.
A much cleaner approach to keeping capacity readily available without relying on Kubernetes scheduling hacks.
-
http://Kubernetes capacity bufferDepending on your level of expertise in this area, you may wonder why Istio’s support for canary deployment is even needed, given that platforms like Kubernetes already provide a way to do version rollout and canary deployment. Problem solved, right? Well, not exactly. Although doing a rollout this way works in simple cases, it’s very limited, especially in large scale cloud environments receiving lots of (and especially varying amounts of) traffic, where autoscaling is needed.
-
https://istio.io/latest/blog/2017/0.1-canary/apiVersion: v1
kind: ConfigMap
metadata:
name: gw-options
data:
horizontalPodAutoscaler: |
spec:
minReplicas: 2
maxReplicas: 2
deployment: |
metadata:
annotations:
additional-annotation: some-value
spec:
replicas: 4
template:
spec:
containers:
- name: istio-proxy
resources:
requests:
cpu: 1234m
service: |
spec:
ports:
- "\$patch": delete
port: 15021
-
https://istio.io/latest/docs/tasks/traffic-management/ingress/gateway-api/#configuring-a-gateway resource.customizations.ignoreDifferences.apps_Deployment: |
jsonPointers:
- /spec/template/metadata/creationTimestamp
resource.customizations.ignoreDifferences.apps_StatefulSet: |
jsonPointers:
- /spec/template/metadata/creationTimestamp
resource.customizations.ignoreDifferences.apps_DaemonSet: |
jsonPointers:
- /spec/template/metadata/creationTimestamp
-
https://github.com/argoproj/argo-cd/issues/25184#issuecomment-3491499482