Autoscaling
This guide is the basic deployment plus per-role autoscaling. The
prefill role sets scaling.mode: External, so the controller creates a
DisaggregatedSetRoleScaler named <ds>-prefill with a /scale subresource.
The bundled HorizontalPodAutoscaler targets that scaler (stable across
rollouts), reads leader pod metrics, and scales between 2 and 8 replicas at 70%
CPU. It needs metrics-server.
External scaling requires spec.slices: 1.
Deploy
export HF_TOKEN=<your-hf-token>
curl https://raw.githubusercontent.com/kubernetes-sigs/lws/refs/heads/main/docs/examples/disaggregatedset/autoscaling/vllm.yaml -s | envsubst | kubectl apply -f -
export HF_TOKEN=<your-hf-token>
curl https://raw.githubusercontent.com/kubernetes-sigs/lws/refs/heads/main/docs/examples/disaggregatedset/autoscaling/sglang.yaml -s | envsubst | kubectl apply -f -
Watch the scaler and HPA:
kubectl get disaggregatedsetrolescaler
kubectl get hpa -w
KEDA can drive the same DisaggregatedSetRoleScaler in place of the HPA: point a
KEDA ScaledObject at the <ds>-prefill scaler. Don’t target the role’s child
LeaderWorkerSet directly — the DisaggregatedSet controller owns each role’s
replica count and overwrites external changes on every reconcile. See the
role scaler concepts.
Feedback
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.