Senior Software Engineer - AI Inference Performance
Core
Optimize LLM and VLM inference performance on NVIDIA GPU-accelerated systems by analyzing workloads, building performance models, and tuning serving software.
Role type
Senior IC software engineer (AI inference performance)
Builds
High-performance inference engines and serving software (e.g., TensorRT-LLM, vLLM, SGLang)
Domain
AI/ML inference, GPU computing, distributed systems
Deliverable
production ML models
Required skills
CUDA programming, Python, Rust, C++, GPU architecture analysis, performance profiling, distributed systems, kernel optimization
Preferred skills
Open-source contributions, AI-agent workflows, published research, multimodal pipeline optimization
Technologies
CUDA, CUTLASS, Triton, Nsight Systems, Nsight Compute, PyTorch, TensorRT-LLM, vLLM, SGLang, NCCL
Responsibilities
Lead end-to-end analysis of LLM/VLM inference processes; Build speed-of-light and roofline models; Profile workloads to eliminate bottlenecks; Tune serving hyperparameters and techniques; Build and optimize performance-critical kernels; Establish repeatable benchmarks and performance regression gates
Seniority
Senior, hands-on IC