4370 links
  • Arnaud's links
  • Home
  • Login
  • RSS Feed
  • ATOM Feed
  • Tag cloud
  • Picture wall
  • Daily
Links per page: 20 50 100
◄Older
page 1 / 219
  • thumbnail
    GitHub - kubernetes-sigs/node-readiness-controller: Declarative node readiness for Kubernetes - gate scheduling with taints until node infrastructure (CNI, GPU, storage, custom checks) is actually ready. · GitHub

    Karpenter: Make Nodes Ready Only After Required DaemonSets Are Ready

    Recommended approach

    Use a startup taint plus a readiness controller. The controller
    should remove the taint only when every DaemonSet explicitly designated
    as required has a Ready pod on that node.

    This keeps Kubernetes’ native NodeReady condition separate from
    workload initialization and prevents ordinary workloads from scheduling
    too early. Avoid relying on a fixed sleep, which can fail when images
    take longer to pull or a DaemonSet is unhealthy.

    The implementation consists of:

    1. A Karpenter NodePool startup taint.
    2. An allowlist of required DaemonSets.
    3. A readiness reporter that checks the pods on each node and publishes
      a custom node condition.
    4. A Node Readiness Controller rule that removes the taint when the
      condition is satisfied.

    1. Add a startup taint to the Karpenter NodePool

    Merge this into your existing NodePool; retain your existing
    nodeClassRef, requirements, limits, and disruption settings.

    apiVersion: karpenter.sh/v1
    kind: NodePool
    metadata:
      name: default
    spec:
      template:
        spec:
          startupTaints:
            - key: readiness.example.com/required-daemonsets
              value: pending
              effect: NoSchedule

    Karpenter uses startup taints for temporary initialization requirements
    and expects another component to remove them.

    Important: Apply this only to the relevant NodePools. Updating a
    NodePool does not necessarily add the startup taint to nodes that
    already exist. Verify how your existing nodes are managed before rolling
    out the change.

    Documentation: Karpenter
    NodePools

    2. Install the Node Readiness Controller

    Use the project’s official release manifests. The version below is an
    example from the installation guide; check the releases
    page
    and current installation instructions before deploying to production.

    VERSION=v0.5.0
    
    kubectl apply -f \
      "https://github.com/kubernetes-sigs/node-readiness-controller/releases/download/${VERSION}/crds.yaml"
    
    kubectl wait \
      --for=condition=Established \
      --timeout=60s \
      crd/nodereadinessrules.readiness.node.x-k8s.io
    
    kubectl apply -f \
      "https://github.com/kubernetes-sigs/node-readiness-controller/releases/download/${VERSION}/install.yaml"

    Project: Kubernetes SIGs Node Readiness
    Controller

    3. Define the readiness rule

    The rule requires a custom node condition named
    readiness.example.com/RequiredDaemonSetsReady to be True. The
    reporter described in the next section is responsible for setting this
    condition.

    apiVersion: readiness.node.x-k8s.io/v1alpha1
    kind: NodeReadinessRule
    metadata:
      name: required-daemonsets-ready
    spec:
      nodeSelector:
        matchLabels:
          readiness.example.com/enabled: "true"
    
      conditions:
        - type: readiness.example.com/RequiredDaemonSetsReady
          requiredStatus: "True"
    
      taint:
        key: readiness.example.com/required-daemonsets
        value: pending
        effect: NoSchedule
    
      enforcementMode: continuous

    The node selector means that only nodes labelled
    readiness.example.com/enabled=true are governed by this rule.

    Continuous enforcement is intended to restore the taint if the required
    readiness condition later becomes false. A NoSchedule taint blocks new
    workloads but does not automatically evict workloads already running on
    the node.

    Check the current API schema and examples in the Node Readiness
    Controller
    documentation
    before applying this manifest.

    4. Define the required DaemonSets

    Use an explicit allowlist of DaemonSets that must be ready on every
    applicable node.

    apiVersion: v1
    kind: ConfigMap
    metadata:
      name: daemonset-readiness-config
      namespace: kube-system
    data:
      required-daemonsets.yaml: |
        required:
          - namespace: kube-system
            name: kube-proxy
          - namespace: kube-system
            name: aws-node
          - namespace: kube-system
            name: ebs-csi-node
          - namespace: monitoring
            name: node-exporter

    Replace these examples with the DaemonSets actually required in your
    cluster. For example, if you use Cilium instead of the AWS VPC CNI,
    replace aws-node with your Cilium DaemonSet.

    5. Implement the readiness reporter

    The Node Readiness Controller does not automatically inspect arbitrary
    DaemonSet pod readiness. A reporter must check the required pods and
    update the custom node condition.

    For each node, the reporter should:

    1. Read the node name from the downward API.
    2. Check each required DaemonSet’s pod on that node.
    3. Confirm that each pod is owned by the expected DaemonSet and has
      Ready=True.
    4. Set readiness.example.com/RequiredDaemonSetsReady=True only when
      every required DaemonSet has a ready pod on that node.
    5. Set the condition to False if a required pod is missing, unready,
      or terminating.

    Treat a missing or unknown readiness signal as not ready.

    Reporter RBAC

    The following is the RBAC portion for a reporter that reads pods and
    DaemonSets and updates node status. It is not a complete deployable
    reporter.

    apiVersion: v1
    kind: ServiceAccount
    metadata:
      name: daemonset-readiness-reporter
      namespace: kube-system
    ---
    apiVersion: rbac.authorization.k8s.io/v1
    kind: ClusterRole
    metadata:
      name: daemonset-readiness-reporter
    rules:
      - apiGroups: [""]
        resources: ["pods", "nodes"]
        verbs: ["get", "list", "watch"]
      - apiGroups: ["apps"]
        resources: ["daemonsets"]
        verbs: ["get", "list", "watch"]
      - apiGroups: [""]
        resources: ["nodes/status"]
        verbs: ["patch", "update"]
    ---
    apiVersion: rbac.authorization.k8s.io/v1
    kind: ClusterRoleBinding
    metadata:
      name: daemonset-readiness-reporter
    roleRef:
      apiGroup: rbac.authorization.k8s.io
      kind: ClusterRole
      name: daemonset-readiness-reporter
    subjects:
      - kind: ServiceAccount
        name: daemonset-readiness-reporter
        namespace: kube-system

    You must still deploy reporter code as a DaemonSet, provide the node
    name through the downward API, and implement the condition-update loop.
    The allowlist is illustrative; do not use it unchanged if those
    DaemonSets are not required in your cluster.

    6. Make required DaemonSets tolerate the startup taint

    Every required DaemonSet must tolerate the custom NoSchedule taint;
    otherwise, its pod may never schedule and the node may remain blocked
    indefinitely.

    Add this toleration to each required DaemonSet’s pod template:

    spec:
      template:
        spec:
          tolerations:
            - key: readiness.example.com/required-daemonsets
              operator: Equal
              value: pending
              effect: NoSchedule

    The readiness reporter must tolerate the same taint. The Node Readiness
    Controller must remain schedulable on an existing healthy node so that
    it can remove taints from newly provisioned nodes.

    7. Validate the rollout

    After deploying the configuration and reporter, inspect the state:

    kubectl get nodepools
    kubectl get nodes
    kubectl get nodereadinessrules
    kubectl get nodes \
      -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints

    Test on a non-production NodePool:

    • Provision a new node.
    • Confirm the startup taint prevents ordinary workloads from
      scheduling.
    • Confirm all required DaemonSet pods become ready.
    • Confirm the reporter sets the custom node condition to True.
    • Confirm the taint is removed.
    • Stop a required DaemonSet pod and verify that continuous enforcement
      restores the taint.

    Important limitations

    • A NoSchedule taint blocks new scheduling; it does not evict
      workloads already running on the node.
    • Evicting existing workloads when a required DaemonSet becomes
      unready needs a separate policy and should be designed carefully to
      avoid unnecessary disruption.
    • Validate the Node Readiness Controller’s current CRD schema, release
      version, and rule semantics against its official documentation
      before applying the sample rule.
    • The reporter is a required implementation component; the YAML above
      provides its configuration and RBAC starting point, not the
      reporter’s executable code.

    References

    • Kubernetes SIGs Node Readiness
      Controller
    • Node Readiness Controller: getting
      started
    • Node Readiness Controller: CNI readiness
      example
    • Karpenter NodePools
      documentation
    October 11, 2026 at 7:45:51 AM GMT+2 * - permalink - archive.org - https://github.com/kubernetes-sigs/node-readiness-controller
    k8s node ready
  • thumbnail
    GitHub - aptible/supercronic: Cron for containers · GitHub
    October 10, 2026 at 8:28:46 PM GMT+2 * - permalink - archive.org - https://github.com/aptible/supercronic
    cron container
  • GitHub - github/spec-kit: 💫 Toolkit to help you get started with SDD or any other process! · GitHub
    October 9, 2026 at 8:15:14 PM GMT+2 * - permalink - archive.org - https://github.com/github/spec-kit
    llm context github spec
  • Amazon ECR expands registry policy to all ECR actions

    aws ecr get-account-setting --name REGISTRY_POLICY_SCOPE
    aws ecr put-account-setting --name REGISTRY_POLICY_SCOPE --value V2

    October 4, 2026 at 4:13:16 PM GMT+2 * - permalink - archive.org - https://aws.amazon.com/about-aws/whats-new/2024/12/amazon-ecr-expands-registry-policy-ecr-actions/
    ecr
  • Blob mounting in Amazon ECR - Amazon ECR
    October 4, 2026 at 4:08:16 PM GMT+2 * - permalink - archive.org - https://docs.aws.amazon.com/AmazonECR/latest/userguide/blob-mounting.html
    ecr
  • Introduction - The Kubebuilder Book
    October 3, 2026 at 1:02:21 PM GMT+2 * - permalink - archive.org - https://book.kubebuilder.io/
    book k8s operator
  • thumbnail
    Owners mourn spoiled food after firmware update bricks Samsung smart fridges - Ars Technica
    September 27, 2026 at 8:08:29 AM GMT+2 * - permalink - archive.org - https://arstechnica.com/gadgets/2026/09/owners-mourn-spoiled-food-after-firmware-update-bricks-samsung-smart-fridges/
    fridge deployment samsung
  • thumbnail
    External Scalers | KEDA
    September 21, 2026 at 7:29:29 AM GMT+2 * - permalink - archive.org - https://keda.sh/docs/2.20/concepts/external-scalers/
    keda external scaler
  • Kubernetes v1.37: Pod Certificates and Cluster Trust Bundles | Kubernetes
    September 6, 2026 at 5:01:34 PM GMT+2 * - permalink - archive.org - https://kubernetes.io/blog/2026/08/28/kubernetes-v1-37-pod-certificates-and-cluster-trust-bundles/
    kubernetes mTLS
  • Note: Karpenter Capacity Buffer

    A long-awaited Karpenter feature is finally here!

    The latest Karpenter release supports Capacity Buffer, eliminating the need for workarounds such as balloon pods to maintain spare node capacity.

    A Capacity Buffer defines virtual placeholder pods that exist only in Karpenter’s scheduling simulation — they are never created as actual Kubernetes pods.

    These virtual pods tell Karpenter to provision nodes with spare capacity. They participate in each scheduling cycle to maintain the buffer and are automatically refilled as real workloads consume the pre-provisioned capacity.

    Benefits over previous workarounds:

    • No balloon pods: Eliminates the need for low-priority placeholder deployments and complex PriorityClass configurations.

    • More efficient scaling: Capacity can be defined using fixed counts or percentages, allowing spare capacity to scale with your workload instead of over-provisioning entire node pools.

    • Automatic replenishment: The buffer is automatically maintained as workloads consume the available capacity.

    A much cleaner approach to keeping capacity readily available without relying on Kubernetes scheduling hacks.

    • 1.14.0 Github release: https://github.com/kubernetes-sigs/karpenter/releases/tag/v1.14.0
    • Documentation: https://karpenter.sh/docs/concepts/capacitybuffers/
    • The outdated baloon pod trick: https://kubernetes.io/docs/tasks/administer-cluster/node-overprovisioning/
    August 19, 2026 at 9:02:52 PM GMT+2 * - permalink - archive.org - http://Kubernetes capacity buffer
    karpenter k8s overprovision buffer
  • Note: public holidays

    fr https://www.service-public.gouv.fr/particuliers/actualites/A18558
    hk https://www.gov.hk/en/about/abouthk/holiday/2026.htm
    uk https://www.gov.uk/bank-holidays

    May 2, 2026 at 8:58:14 PM GMT+2 * - permalink - archive.org - https://links.infomee.fr/shaare/oSVzoQ
    public bank holidays
  • thumbnail
    Notion skills
    March 22, 2026 at 1:23:58 PM GMT+1 * - permalink - archive.org - https://www.notion.so/notiondevs/Notion-Skills-for-Claude-28da4445d27180c7af1df7d8615723d0
    claude notion skill
  • https://github.com/anthropics/skills/blob/main/skills/doc-coauthoring/SKILL.md
    March 22, 2026 at 1:21:32 PM GMT+1 - permalink - archive.org - https://github.com/anthropics/skills/blob/main/skills/doc-coauthoring/SKILL.md
    claude skill doc
  • thumbnail
    GitHub - sirmalloc/ccstatusline: 🚀 Beautiful highly customizable statusline for Claude Code CLI with powerline support, themes, and more. · GitHub

    To help generate claude status line

    March 22, 2026 at 11:44:07 AM GMT+1 * - permalink - archive.org - https://github.com/sirmalloc/ccstatusline
    claude
  • GitHub - ldayton/Dippy: 🐤 Less permission fatigue, more momentum. Dippy knows what’s safe to run and keeps Claude on track when plans change. · GitHub
    March 21, 2026 at 5:07:38 PM GMT+1 - permalink - archive.org - https://github.com/ldayton/Dippy
    claude permissions
  • thumbnail
    Istio / Canary Deployments using Istio

    Depending on your level of expertise in this area, you may wonder why Istio’s support for canary deployment is even needed, given that platforms like Kubernetes already provide a way to do version rollout and canary deployment. Problem solved, right? Well, not exactly. Although doing a rollout this way works in simple cases, it’s very limited, especially in large scale cloud environments receiving lots of (and especially varying amounts of) traffic, where autoscaling is needed.

    March 8, 2026 at 8:41:32 PM GMT+1 * - permalink - archive.org - https://istio.io/latest/blog/2017/0.1-canary/
    istio canary
  • thumbnail
    Istio / Kubernetes Gateway API
    apiVersion: v1
    kind: ConfigMap
    metadata:
      name: gw-options
    data:
      horizontalPodAutoscaler: |
        spec:
          minReplicas: 2
          maxReplicas: 2
    
      deployment: |
        metadata:
          annotations:
            additional-annotation: some-value
        spec:
          replicas: 4
          template:
            spec:
              containers:
              - name: istio-proxy
                resources:
                  requests:
                    cpu: 1234m
    
      service: |
        spec:
          ports:
          - "\$patch": delete
            port: 15021
    March 8, 2026 at 7:57:42 PM GMT+1 * - permalink - archive.org - https://istio.io/latest/docs/tasks/traffic-management/ingress/gateway-api/#configuring-a-gateway
    istio
  • thumbnail
    Istio / Istio Standard Metrics
    March 1, 2026 at 12:07:55 PM GMT+1 - permalink - archive.org - https://istio.io/latest/docs/reference/config/metrics/
    istio metrics
  • thumbnail
    Library Charts | Helm
    • https://github.com/ksemele/demo-helm-library
    • https://ksemele.medium.com/how-to-migrate-from-helm-monorepo-to-versioned-charts-66dfe5db321b
    February 28, 2026 at 9:51:39 AM GMT+1 - permalink - archive.org - https://helm.sh/docs/topics/library_charts/
    helm library
  • Server-Side Diff shows diff on deployment.spec.template.metadata.creationTimestamp in v3.2.0 · Issue #25184 · argoproj/argo-cd · GitHub
      resource.customizations.ignoreDifferences.apps_Deployment: |
        jsonPointers:
          - /spec/template/metadata/creationTimestamp
      resource.customizations.ignoreDifferences.apps_StatefulSet: |
        jsonPointers:
          - /spec/template/metadata/creationTimestamp
      resource.customizations.ignoreDifferences.apps_DaemonSet: |
        jsonPointers:
          - /spec/template/metadata/creationTimestamp
    
    February 27, 2026 at 6:13:05 AM GMT+1 - permalink - archive.org - https://github.com/argoproj/argo-cd/issues/25184#issuecomment-3491499482
    argocd
Links per page: 20 50 100
◄Older
page 1 / 219
Shaarli - The personal, minimalist, super fast, database-free, bookmarking service by the Shaarli community - Help/documentation