CareerPlanSign in

Sr Staff Site Reliability Engineer, AI Infrastructure

Santa Clara💼 Full-time🗓 2026-09-29 → 2026-10-01

Core

Own reliability, automation, and observability for colocation, on-prem GPU clusters, and cloud environments supporting AI hardware/software deployment.

Role type

Senior Staff Site Reliability Engineer (AI Infrastructure)

Builds

Production infrastructure for AI workloads, HPC clusters, and customer-facing platform services

Domain

AI Infrastructure / High-Performance Computing / Cloud & On-Prem

Deliverable

infrastructure

Required skills

Linux systems administration, Infrastructure as Code (Terraform/Ansible), Kubernetes operations, Observability (Prometheus/Grafana/DataDog/Splunk), Python/Bash scripting, Incident response and RCA, Capacity planning, Hardware lifecycle management

Preferred skills

Customer-facing infrastructure operations, Hybrid cloud environments, AIOps platforms, HPC job schedulers (Slurm/LSF), High-speed interconnects (InfiniBand/RoCE/NVLink), Large-scale automation tooling

Responsibilities

Own reliability across server fleets and cloud environments; Perform hands-on server provisioning, networking, and hardware troubleshooting; Lead capacity planning and cloud spend tracking; Drive infrastructure changes via IaC; Build automation for fleet health and self-service tooling; Design monitoring and alerting workflows; Participate in on-call rotation and produce RCAs for critical incidents.

Seniority

Senior Staff, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 901,000+ jobs from 20+ sources.