Software Engineer, DGX Cloud AI Infrastructure - New College Grad 2026
Core
Design, implement, and optimize distributed training and inference workloads for large-scale LLMs on NVIDIA GPU platforms.
Role type
Senior IC software engineer (distributed AI systems)
Builds
Benchmarking tooling, automation, debugging workflows, and resilience systems for multi-GPU/multi-node clusters
Domain
Generative AI, HPC, distributed computing, GPU infrastructure
Deliverable
production ML models
Required skills
Python, C/C++, CUDA, PyTorch, NeMo, TensorRT-LLM, distributed systems debugging, cluster operations, root-cause analysis
Preferred skills
NCCL, RDMA stack (IB verbs, UCX, libfabric), InfiniBand/RoCE congestion debugging, MLPerf, performance jitter diagnosis, fault-detection systems
Technologies
PyTorch, NeMo, Megatron, TensorRT-LLM, CUDA, NCCL, InfiniBand, RoCE, UCX, libfabric
Responsibilities
Bring up and debug large-scale AI clusters; tune runtime settings and deployment configurations; build repeatable benchmark suites and qualification workflows; contribute to failure-attribution tooling; perform root-cause analysis of distributed environment failures
Seniority
Junior (New College Grad 2026), hands-on IC