Member of Technical Staff, Cloud Orchestration
Core
Design and build operational systems for cluster management, deployment automation, and production monitoring to keep vLLM running reliably at massive scale.
Role type
Senior IC cloud orchestration engineer (GPU inference)
Builds
vLLM inference clusters and deployment pipelines
Domain
AI inference infrastructure / GPU cluster management
Deliverable
infrastructure
Required skills
Kubernetes at scale, custom Kubernetes operators, Python/Rust/Go, infrastructure-as-code (Terraform, Helm), GPU cluster management, cloud platform expertise (AWS/GCP/Azure)
Preferred skills
ML-specific orchestration (Ray, Slurm), GPU scheduling, multi-tenancy, vLLM deployment patterns, operational reliability for ML systems
Technologies
Kubernetes, Terraform, Helm, Ray, Slurm, vLLM, AWS, GCP, Azure
Responsibilities
Design systems for cluster management and deployment automation; ensure vLLM deployments are observable, debuggable, and recoverable; manage GPU clusters and debug hardware issues; work across cloud platforms and on-premise infrastructure
Seniority
Senior, hands-on IC