CareerPlanSign in

Principal Software Engineer, Distributed Systems Engineer - DGX Cloud

US, NC, Durham💼 Full-time💰 $272,000–$272,000🗓 2026-09-29 → 2026-10-01

Core

Designing and scaling production AI infrastructure for large GPU clusters using Kubernetes to support diverse AI workloads.

Role type

Principal Software Engineer, Distributed Systems

Builds

Scalable GPU clusters and custom Kubernetes scheduling software for AI applications

Domain

AI computing, GPU infrastructure, Cloud-native systems

Deliverable

production ML models | infrastructure

Required skills

Kubernetes API development, Cluster operations, GPU resource scheduling, Systems programming (Go/Python), Data structures and algorithms, Incident management, Large-scale distributed systems design

Preferred skills

Cluster management systems (Slurm, Bright Cluster Manager), Independent cloud-agnostic system management, Operational excellence in AI infrastructure

Responsibilities

Develop custom software for GPU resource scheduling on Kubernetes, Implement monitoring and health management for GPU assets, Evaluate system failures and improve services via incident management processes

Seniority

Principal, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 902,000+ jobs from 20+ sources.