CareerPlanSign in

Member of Technical Staff, Kernel Engineering

San Francisco💼 Full-time💰 $200,000–$200,000🗓 2026-09-30 → 2026-10-01

Core

Write CUDA kernels and low-level optimizations to maximize performance of the vLLM AI inference engine on modern accelerators.

Role type

Senior IC performance engineer (GPU kernel development)

Builds

High-performance AI inference engine (vLLM) for diverse hardware accelerators

Domain

AI inference, GPU architecture, high-performance computing

Deliverable

production ML models

Required skills

CUDA kernel development, GPU architecture knowledge, C++, Python, performance profiling, benchmarking

Preferred skills

ML-specific kernel optimization, quantization techniques, multi-platform accelerator experience, compiler technologies

Technologies

CUDA, CuTeDSL, Triton, TileLang, Pallas, Nsight, rocprof, LLVM, MLIR, XLA

Responsibilities

Write and optimize kernels for NVIDIA GPUs and emerging silicon, integrate with hardware vendor teams, profile and benchmark performance

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 902,000+ jobs from 20+ sources.