CareerPlanSign in

Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Seattle, United States of America💼 Full-time🗓 2026-08-13 → 2026-09-28

Core

Building and optimizing high-performance, low-latency inference systems for large foundation models (language, vision, speech) serving billions of queries across Apple products.

Role type

Senior IC machine learning engineer (foundation model inference)

Builds

Production-grade inference systems and tooling for planetary-scale deployment

Domain

Cloud infrastructure + Large Language Models (LLMs) + Multimodal AI

Deliverable

production ML models

Required skills

LLM inference stack optimization, GPU/TPU programming, PyTorch/JAX/TensorFlow, high-throughput distributed services, cloud platform deployment (Kubernetes/Docker)

Preferred skills

Go/Python production systems, deep learning architectures (Transformers, multimodal), inference optimization frameworks (TensorRT-LLM, vLLM, SGLang, TGI, Triton), custom CUDA kernel development

Technologies

PyTorch, JAX, TensorFlow, Kubernetes, Docker, CUDA C++, OpenAI Triton, TensorRT-LLM, vLLM, SGLang, TGI, Nvidia Triton Server

Responsibilities

Optimize inference for latest model architectures across language, vision, and speech; Design and ship production-grade inference systems; Build profiling tools and simulators for performance bottlenecks; Drive technical decisions on high-throughput, low-latency serving; Mentor and grow engineers

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via apple · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.