Software Engineer, CUDA Deep Learning Systems
Core
Designing and optimizing high-performance CUDA kernels and distributed computing systems for next-generation deep learning models and AI workloads.
Role type
Senior IC CUDA Deep Learning Systems Engineer
Builds
Custom CUDA kernels, distributed training/inference systems, and runtime profiling tools for AI accelerators
Domain
AI Infrastructure / High-Performance Computing / GPU Systems
Deliverable
production ML models
Required skills
C++, Python, CUDA programming, deep learning fundamentals (transformers), distributed computing, systems programming, computer architecture, performance optimization, generative AI model profiling
Preferred skills
PyTorch/JAX/TensorRT/vLLM internals, NCCL/MPI/UCX, low-precision arithmetic (FP8/INT8), deep learning compilers (Triton/XLA), agentic AI systems design
Technologies
CUDA, C++, Python, PyTorch, JAX, TensorRT, vLLM, NCCL, MPI, UCX, Triton, XLA
Responsibilities
Research and prototype systems optimizations for deep learning models; Architect and optimize distributed computing systems from single node to cluster-scale; Design and implement custom high-performance CUDA kernels; Analyze hardware-software interactions to resolve performance bottlenecks; Collaborate with researchers and architects to improve compute utilization and network efficiency; Develop exploratory tools for profiling and accelerating new deep learning paradigms
Seniority
Senior, hands-on IC