AI Infrastructure System Engineer
Core
Design and build automated systems to provision, validate, deploy, upgrade, repair, and retire large-scale GPU clusters for AI training and inference workloads.
Role type
Senior IC AI Infrastructure System Engineer
Builds
Automated fleet management platforms, AI infrastructure agents, and validation frameworks for GPU clusters
Domain
AI Infrastructure / Distributed Systems / GPU Computing
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Rust, Linux, Kubernetes, Terraform, Ansible, distributed systems, GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch, InfiniBand, RoCE, bare-metal provisioning, predictive failure detection
Preferred skills
AI agents, autonomous infrastructure operations, distributed storage systems
Technologies
Python, Go, Rust, Kubernetes, Terraform, Ansible, CUDA, NCCL, NVLink, NVSwitch, InfiniBand, RoCE
Responsibilities
Design fleet automation systems for GPU cluster lifecycle management; Develop AI infrastructure agents for incident triage and remediation; Build fleet intelligence platforms for hardware and workload monitoring; Create automated validation frameworks for GPUs and networking fabrics; Develop internal platforms for programmable infrastructure management; Collaborate with hardware, networking, and AI teams on scalable solutions
Seniority
Senior, hands-on IC