Failure Handling and Restart Policies

Learn how LeaderWorkerSet handles pod and node failures with configurable restart policies.

LeaderWorkerSet provides configurable failure handling for pod groups, ensuring that pod and node failures in distributed workloads are handled consistently according to the coupling requirements of the application.

Configure the failure and restart behavior via .spec.leaderWorkerTemplate.restartPolicy:

RecreateGroupOnPodRestart (Default)

When any pod in a group fails or restarts, the entire replica group (leader + all workers) is deleted and recreated.

  • Pod Failures: If a single container or pod fails or restarts, all other pods in the group are terminated and recreated simultaneously to ensure all processes restart fresh and re-initialize collective communication or distributed caches cleanly.
  • Node Failures: When a node hosting any pod in the replica fails or becomes unreachable, the entire replica group is deleted and recreated on healthy nodes, respecting topology placement constraints.
  • Primary Use Case: Tightly coupled multi-host distributed inference and training (e.g., tensor-parallel or pipeline-parallel models) where a single pod or node failure breaks collective communication.
apiVersion: leaderworkerset.x-k8s.io/v1
kind: LeaderWorkerSet
metadata:
  name: leaderworkerset-sample
spec:
  replicas: 2
  leaderWorkerTemplate:
    restartPolicy: RecreateGroupOnPodRestart
    size: 4
    workerTemplate:
      spec:
        containers:
        - name: worker
          image: worker-image:latest

None

Only the failed pod is restarted or rescheduled. Other pods in the group continue running without interruption.

  • Pod Failures: If an individual pod or container fails, only that specific pod is restarted by Kubernetes.
  • Node Failures: When a node fails, only the pods residing on that failed node are rescheduled. Other pods in the replica remain running on their existing nodes.
  • Primary Use Case: Loosely coupled workers or workloads with application-level fault tolerance where individual pods can reconnect or recover independently.
apiVersion: leaderworkerset.x-k8s.io/v1
kind: LeaderWorkerSet
metadata:
  name: leaderworkerset-sample
spec:
  replicas: 2
  leaderWorkerTemplate:
    restartPolicy: None
    size: 4
    workerTemplate:
      spec:
        containers:
        - name: worker
          image: worker-image:latest

RecreateGroupAfterStart

When any pod in a group fails, the entire group is recreated if and only if there are no pods currently pending in the group. If any pod in the replica is still in the Pending phase (e.g., during image pulls or initial scheduling), the controller skips the failure event without triggering a group-wide recreation.

  • Pod Failures: Recreates the entire group if a pod fails after all pods in the replica have started (no pods are Pending). If any pod in the replica is Pending, the failure event is skipped, allowing Kubernetes to handle pod restarts individually and preventing restart cascades during rollout.
  • Node Failures: If a node fails after all pods in the replica have started, the entire replica group is deleted and recreated on healthy nodes. If the failure occurs while any pod in the replica is Pending, group recreation is not triggered.
  • Primary Use Case: Workloads with large container images or long startup times where you want strict collective restart semantics in production once running, but want to prevent recreation loops during the initial rollout.
apiVersion: leaderworkerset.x-k8s.io/v1
kind: LeaderWorkerSet
metadata:
  name: leaderworkerset-sample
spec:
  replicas: 2
  leaderWorkerTemplate:
    restartPolicy: RecreateGroupAfterStart
    size: 4
    workerTemplate:
      spec:
        containers:
        - name: worker
          image: worker-image:latest

Limit Automatic Group Recreation

For RecreateGroupOnPodRestart and RecreateGroupAfterStart, set maxGroupRestarts to limit how many times LWS can automatically recreate each replica group. When the field is unset, group recreation remains unlimited. A value of 0 disables automatic group recreation on the first qualifying failure.

apiVersion: leaderworkerset.x-k8s.io/v1
kind: LeaderWorkerSet
metadata:
  name: leaderworkerset-sample
spec:
  replicas: 2
  leaderWorkerTemplate:
    restartPolicy: RecreateGroupOnPodRestart
    maxGroupRestarts: 3
    size: 4
    workerTemplate:
      spec:
        containers:
        - name: worker
          image: worker-image:latest

The budget is tracked independently for each replica group and Pod template revision. It is consumed only when LWS initiates a group recreation. The field is supported with both Ordinal and Hash group identity modes, and is not supported with restartPolicy: None. In Hash mode, a recreated group receives a new group key, and LWS transfers the recreating group’s restart count when admitting the replacement leader so the budget persists across replacements.

When a group exhausts its budget, LWS:

  1. Terminates the leader and worker Pods to release their scheduled resources.
  2. Retains Pod API objects that can still receive cleanup finalizers and stops automatic recreation of that group (in Hash mode, the retained leader holds back its gated replacement leader under both PostTermination and Immediate replacement policies). A Pod that was already deleting may disappear because Kubernetes does not allow adding a finalizer at that point.
  3. Sets Degraded=True with reason ReplicaRestartBudgetExceeded. Other replica groups continue running.

Retaining a Pod API object preserves its status, but does not guarantee that kubectl logs remains available after the container runtime removes the terminated container. Use an external logging system when logs must survive group termination.

Recover an Exhausted Group

After fixing the underlying problem, explicitly recover one exhausted group by annotating its retained leader Pod:

kubectl annotate pod <leader-pod-name> leaderworkerset.sigs.k8s.io/recover=true

LWS then clears that group’s count, removes the cleanup finalizers, and allows a replacement group to start (created by the StatefulSet in Ordinal mode, or admitted from the gated replacement leader in Hash mode) with a fresh budget. Editing or unsetting maxGroupRestarts, or deleting retained Pods, does not recover an exhausted group. LWS deletion, scale-down, and selecting the group for replacement during a rollout remove retained objects as normal lifecycle cleanup rather than starting recovery.

Feedback

Was this page helpful?

Last modified October 1, 2026: Updated MaxGroupRestart docs (#1116) (6827217)