CareerPlanSign in

Member of Technical Staff, TPU Performance Engineering

San Francisco💼 Full-time💰 $200,000–$200,000🗓 2026-09-30 → 2026-10-02

Core

Build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure to make vLLM a first-class inference engine on Google TPUs.

Role type

Senior IC TPU performance engineer (inference systems)

Builds

Production-relevant model serving on TPU hardware with clear correctness, latency, and throughput benchmarks

Domain

AI inference systems, TPU hardware, compilers, and ML kernels

Deliverable

production ML models

Required skills

TPU workload optimization, JAX, XLA, Pallas, ML kernel optimization, performance profiling, benchmarking, TPU execution and memory behavior understanding

Preferred skills

vLLM, SGLang, TensorRT-LLM, XLA-based serving, compiler technologies (MLIR, LLVM), quantization methods (INT8, FP8, mixed precision)

Technologies

JAX, XLA, Pallas, vLLM, SGLang, TensorRT-LLM, MLIR, LLVM

Responsibilities

Build and optimize TPU backends and compiler integrations; optimize ML kernels and inference paths (attention, GEMM, sampling, KV cache); develop benchmarking infrastructure; profile and measure performance to guide optimization work

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 928,000+ jobs from 20+ sources.