Multi-slice
Fan a disaggregated role set out into independent slices with spec.slices.
This guide is the basic deployment plus slice fan-out. spec.slices: 2 replicates the whole prefill/decode role set into two independent slices, so
each slice is a self-contained prefill+decode unit. Combine this with
topology-aware scheduling to co-locate each
slice’s roles and spread slices across domains.
Slice fan-out and external role autoscaling are mutually exclusive: external
scaling requires spec.slices: 1.
Deploy
export HF_TOKEN=<your-hf-token>
curl https://raw.githubusercontent.com/kubernetes-sigs/lws/refs/heads/main/docs/examples/disaggregatedset/multi-slice/vllm.yaml -s | envsubst | kubectl apply -f -
export HF_TOKEN=<your-hf-token>
curl https://raw.githubusercontent.com/kubernetes-sigs/lws/refs/heads/main/docs/examples/disaggregatedset/multi-slice/sglang.yaml -s | envsubst | kubectl apply -f -
Verify the slices:
kubectl get leaderworkersets
See the slices concepts for details.
Feedback
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.