Basic
Minimal multi-node inference on LeaderWorkerSet with vLLM and SGLang.
These guides deploy a distributed, multi-node inference service with
LeaderWorkerSet, spreading tensor and pipeline parallelism across the leader and
worker pods. Each guide isolates one feature and ships both a vllm.yaml and a
sglang.yaml.
exclusive-topology.Minimal multi-node inference on LeaderWorkerSet with vLLM and SGLang.
Scale LeaderWorkerSet replica groups with a HorizontalPodAutoscaler.
Pin each replica group to a single topology domain with exclusive-topology.
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.