Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Core
Define customer workflows, prototype stack enhancements, and debug complex scheduling and runtime issues in multi-rack, multi-tenant AI/ML datacenters.
Role type
Senior IC cloud-native software engineer (Kubernetes/Slurm/GPU)
Builds
Cloud-native stack for NVIDIA GB200/GB300 GPU datacenter products
Domain
AI/ML infrastructure, High-Performance Computing, Cloud Services
Deliverable
production ML models
Required skills
Kubernetes internals (scheduler, CRI/CNI/CSI, operators), Slurm (federation, plugins), GPU integration (Blackwell/GB200/GB300), distributed systems debugging, RDMA/RoCE networking, Go/Rust/C/C++/Python, Infrastructure-as-Code (Helm/Ansible/Terraform), CI/CD pipelines
Preferred skills
CUDA, deep learning workloads, upstream contributions to Kubernetes/Slurm/Volcano
Responsibilities
Debug multi-rack cluster scheduler behavior and container runtime issues, prototype feature extensions for Kubernetes operators and Slurm plugins, drive joint architecture reviews and convert findings into RFCs, create reproducible testbeds and automate validation suites, deliver technical documentation and present at customer events, collaborate with sales and solution architect teams
Seniority
Senior, hands-on IC