Deep Learning Performance Software Intern - 2027
Core
Developing GPU-accelerated deep learning software and highly optimized deep learning kernels using tile-based GPU programming models.
Role type
Deep Learning Performance Software Engineering Intern
Builds
SKILL, Wiki, agent harness, TileGym, Triton CUDA TileIR backend, CUDA Tile
Domain
Deep Learning, GPU Computing, High-Performance Computing
Deliverable
production ML models
Required skills
C/C++ programming, software design, agentic systems (LLM APIs, prompting, tool use, agent loops, planning, reasoning, memory, RAG, skills, MCP, sub-agents, multi-agent architectures, context engineering, harness engineering), performance modeling, profiling, debugging, code optimization, architectural knowledge of CPU and GPU
Preferred skills
Python, MLIR, GPU programming (CUDA or OpenCL), masters or doctoral degree
Technologies
CUDA, Triton, MLIR, C/C++, Python, LLM APIs, RAG, MCP
Responsibilities
Creating and maintaining SKILL, Wiki, and agent harness; Developing TileGym, Triton CUDA TileIR backend, and CUDA Tile; Developing highly optimized deep learning kernels through tile-based GPU programming model; Performing end-to-end performance optimization through tile-based GPU programming model; Conducting performance optimization, analysis, and tuning
