Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Core
Building and optimizing high-performance, low-latency inference systems for large foundation models (language, vision, speech) serving billions of queries across Apple products.
Role type
Senior IC machine learning engineer (foundation model inference)
Builds
Production-grade inference systems and tooling for planetary-scale deployment
Domain
Cloud infrastructure + Large Language Models (LLMs) + Multimodal AI
Deliverable
production ML models
Required skills
LLM inference stack optimization, GPU/TPU programming, PyTorch/JAX/TensorFlow, high-throughput distributed services, cloud platform deployment (Kubernetes/Docker)
Preferred skills
Go/Python production systems, deep learning architectures (Transformers, multimodal), inference optimization frameworks (TensorRT-LLM, vLLM, SGLang, TGI, Triton), custom CUDA kernel development
Technologies
PyTorch, JAX, TensorFlow, Kubernetes, Docker, CUDA C++, OpenAI Triton, TensorRT-LLM, vLLM, SGLang, TGI, Nvidia Triton Server
Responsibilities
Optimize inference for latest model architectures across language, vision, and speech; Design and ship production-grade inference systems; Build profiling tools and simulators for performance bottlenecks; Drive technical decisions on high-throughput, low-latency serving; Mentor and grow engineers
Seniority
Senior, hands-on IC with mentorship responsibilities