Sr Staff Site Reliability Engineer, AI Infrastructure
Core
Own reliability, automation, and observability for colocation, on-prem GPU clusters, and cloud environments supporting AI hardware/software deployment.
Role type
Senior Staff Site Reliability Engineer (AI Infrastructure)
Builds
Production infrastructure for AI workloads, HPC clusters, and customer-facing platform services
Domain
AI Infrastructure / High-Performance Computing / Cloud & On-Prem
Deliverable
infrastructure
Required skills
Linux systems administration, Infrastructure as Code (Terraform/Ansible), Kubernetes operations, Observability (Prometheus/Grafana/DataDog/Splunk), Python/Bash scripting, Incident response and RCA, Capacity planning, Hardware lifecycle management
Preferred skills
Customer-facing infrastructure operations, Hybrid cloud environments, AIOps platforms, HPC job schedulers (Slurm/LSF), High-speed interconnects (InfiniBand/RoCE/NVLink), Large-scale automation tooling
Responsibilities
Own reliability across server fleets and cloud environments; Perform hands-on server provisioning, networking, and hardware troubleshooting; Lead capacity planning and cloud spend tracking; Drive infrastructure changes via IaC; Build automation for fleet health and self-service tooling; Design monitoring and alerting workflows; Participate in on-call rotation and produce RCAs for critical incidents.
Seniority
Senior Staff, hands-on IC