CareerPlanSign in

AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure

San Francisco Bay Area, United States of America💼 Full-time🗓 2026-08-31 → 2026-09-28

Core

Drive performance optimization for large-scale foundation model training on TPUs, focusing on efficiency, throughput, and scalability.

Role type

Staff ML Infrastructure Engineer (Pre-training)

Builds

High-performance TPU kernels and distributed training systems for foundation models

Domain

AI/ML Infrastructure, High-Performance Computing (HPC)

Deliverable

production ML models

Required skills

Python, distributed systems, parallel computing, performance optimization, TPU architecture, JAX, XLA, collective communication, kernel development (Pallas/Triton/CUDA)

Preferred skills

Advanced degree, GPU accelerator experience, PyTorch, large-scale foundation model training optimization

Technologies

TPU, JAX, XLA, Pallas, Triton, CUDA, PyTorch, ICI/Fabric

Responsibilities

Profile and optimize JAX/XLA workloads across compute, memory, and communication; Develop and optimize high-performance TPU kernels for attention and MoE; Optimize distributed training techniques and sharding strategies; Research and implement new techniques across the JAX, XLA, and TPU stack; Develop performance profiling, benchmarking, and automated tuning capabilities; Lead complex technical projects and mentor engineers

Seniority

Staff, hands-on IC with mentorship

Sourced via apple · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.