CareerPlanSign in

AI Infrastructure System Engineer

India🌐 Remote💼 Full-time🗓 2026-09-25 → 2026-09-29

Core

Design and build automated systems to provision, validate, deploy, upgrade, repair, and retire large-scale GPU clusters for AI training and inference workloads.

Role type

Senior IC AI Infrastructure System Engineer

Builds

Automated fleet management platforms, AI infrastructure agents, and validation frameworks for GPU clusters

Domain

AI Infrastructure / Distributed Systems / GPU Computing

Deliverable

production ML models | infrastructure

Required skills

Python, Go, Rust, Linux, Kubernetes, Terraform, Ansible, distributed systems, GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch, InfiniBand, RoCE, bare-metal provisioning, predictive failure detection

Preferred skills

AI agents, autonomous infrastructure operations, distributed storage systems

Technologies

Python, Go, Rust, Kubernetes, Terraform, Ansible, CUDA, NCCL, NVLink, NVSwitch, InfiniBand, RoCE

Responsibilities

Design fleet automation systems for GPU cluster lifecycle management; Develop AI infrastructure agents for incident triage and remediation; Build fleet intelligence platforms for hardware and workload monitoring; Create automated validation frameworks for GPUs and networking fabrics; Develop internal platforms for programmable infrastructure management; Collaborate with hardware, networking, and AI teams on scalable solutions

Seniority

Senior, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 872,000+ jobs from 20+ sources.