CareerPlanSign in

Senior Software Engineer - AI Inference Performance

US, CA, Santa Clara💼 Full-time💰 $184,000–$184,000🗓 2026-09-28 → 2026-09-29

Core

Optimize LLM and VLM inference performance on NVIDIA GPU-accelerated systems by analyzing workloads, building performance models, and tuning serving software.

Role type

Senior IC software engineer (AI inference performance)

Builds

High-performance inference engines and serving software (e.g., TensorRT-LLM, vLLM, SGLang)

Domain

AI/ML inference, GPU computing, distributed systems

Deliverable

production ML models

Required skills

CUDA programming, Python, Rust, C++, GPU architecture analysis, performance profiling, distributed systems, kernel optimization

Preferred skills

Open-source contributions, AI-agent workflows, published research, multimodal pipeline optimization

Technologies

CUDA, CUTLASS, Triton, Nsight Systems, Nsight Compute, PyTorch, TensorRT-LLM, vLLM, SGLang, NCCL

Responsibilities

Lead end-to-end analysis of LLM/VLM inference processes; Build speed-of-light and roofline models; Profile workloads to eliminate bottlenecks; Tune serving hyperparameters and techniques; Build and optimize performance-critical kernels; Establish repeatable benchmarks and performance regression gates

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.