Autoscaling
Scale LeaderWorkerSet replica groups with a HorizontalPodAutoscaler.
This guide is the basic deployment plus a HorizontalPodAutoscaler.
The HPA scales the number of replica groups through the LWS scale
subresource (it monitors leader pods only), between minReplicas: 2 and
maxReplicas: 5 at 50% CPU utilization. It needs
metrics-server.
Deploy
export HF_TOKEN=<your-hf-token>
curl https://raw.githubusercontent.com/kubernetes-sigs/lws/refs/heads/main/docs/examples/leaderworkerset/autoscaling/vllm.yaml -s | envsubst | kubectl apply -f -
export HF_TOKEN=<your-hf-token>
curl https://raw.githubusercontent.com/kubernetes-sigs/lws/refs/heads/main/docs/examples/leaderworkerset/autoscaling/sglang.yaml -s | envsubst | kubectl apply -f -
Watch the HPA react to load:
kubectl get hpa -w
See basic for how to reach the service once pods are running.
Feedback
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.