Software Engineer Intern, AI and DL Kernel Libraries - 2027
Core
Develop foundational software and low-level kernels for NVIDIA's AI platform, focusing on deep learning primitives, runtime systems, and performance-critical GPU software for large language models and generative AI.
Role type
Software Engineer Intern (AI/DL Kernel Libraries)
Builds
NVIDIA's AI software stack including cuDNN, FlashInfer, and optimized support for LLM inference workloads.
Domain
AI/Deep Learning Systems, GPU Computing, High-Performance Computing
Deliverable
production ML models
Required skills
C/C++, Python, CUDA development, deep learning frameworks (PyTorch, JAX, TensorFlow, ONNX), linear algebra, performance analysis, code optimization
Preferred skills
Inference engines (vLLM, SGLang, TensorRT-LLM), domain-specific compilers, MLIR/Apache TVM, GPU performance modeling, open-source contributions
Technologies
CUDA, cuDNN, FlashInfer, PyTorch, JAX, TensorFlow, ONNX, vLLM, SGLang, TensorRT-LLM, MLIR, Apache TVM
Responsibilities
Design and optimize kernels for high-impact AI workloads; build extensible software abstractions for deep learning libraries and runtime systems; contribute to just-in-time compilation and code generation; analyze workload performance and tune software; collaborate on open-source ecosystem integrations.
Seniority
Intern