Cloud Systems Engineer
Core
Operate and maintain large-scale GPU-accelerated compute infrastructure for AI, machine learning, and high-performance computing (HPC) workloads.
Role type
Senior Cloud Systems Engineer (Infrastructure)
Builds
GPU-accelerated compute clusters for AI model training, inference, and data processing
Domain
Cloud Infrastructure / AI / HPC
Deliverable
infrastructure
Required skills
Linux administration, GPU hardware troubleshooting, firmware lifecycle management, Bash/Python/PowerShell scripting, storage and networking fundamentals, infrastructure monitoring (Grafana), root-cause analysis
Preferred skills
Enterprise server infrastructure support, datacenter operations, InfiniBand networking, automation development
Responsibilities
Deploy and maintain GPU-accelerated compute infrastructure; manage OS, firmware, BIOS, and driver updates; monitor system health and performance; troubleshoot hardware and OS issues; develop operational runbooks; perform rack-and-stack deployments and hardware replacements; support high-performance networking environments.
Seniority
Mid-Senior, hands-on IC