Member of Technical Staff, Cluster Administration
Core
Own and operate high-performance GPU compute infrastructure to keep engineering teams productive.
Role type
Senior IC cluster administration engineer (GPU/HPC)
Builds
High-performance GPU and HPC clusters across neo-cloud and dedicated providers
Domain
AI inference infrastructure / High-Performance Computing
Deliverable
infrastructure
Required skills
Linux systems administration, GPU server operations, cluster scheduling (SLURM/Kubernetes), infrastructure automation (Bash/Python/Ansible/Terraform), incident response, hardware diagnostics
Preferred skills
Multi-provider GPU operations (Lambda/CoreWeave/etc.), cluster utilization optimization, high-performance GPU networking (InfiniBand/RoCE/NVLink), HPC storage systems (NFS/Lustre/Ceph), infrastructure security hygiene
Responsibilities
Monitor cluster health and GPU availability, manage GPU drivers and hardware diagnostics, handle urgent infrastructure incidents, automate operational workflows, standardize provisioning and operating patterns across providers
Seniority
Senior, hands-on IC