Staff ML Infrastructure Engineer
Core
Architecting and delivering large-scale distributed data and ML infrastructure platforms that power Apple's generative AI model training and inference fleets.
Role type
Staff ML Infrastructure Engineer (Platform Architecture)
Builds
High-throughput data ingestion, versioning, lineage, and delivery systems for petabyte-scale model training on GPU/TPU fleets.
Domain
Generative AI, Distributed Data Systems, ML Infrastructure
Deliverable
production ML models
Required skills
Large-scale distributed data systems, Python, Rust, C++/Go, I/O-bound performance engineering, Columnar/Lakehouse formats (Parquet, Iceberg, Delta), End-to-end ML workflow, Generative AI techniques (Transformers, RAG), System architecture, Technical leadership
Preferred skills
PyTorch/JAX/TensorFlow data-loading layers, Ray Data/NVIDIA DALI, GPU/TPU fleet saturation, Data lineage/governance (DataHub, Unity Catalog), Spark/Daft/Polars internals, Kubernetes
Technologies
Python, Rust, C++, Go, Parquet, Iceberg, Delta, Lance, PyTorch, JAX, TensorFlow, Ray Data, NVIDIA DALI, WebDataset, Spark, Kubernetes
Responsibilities
Define architecture for petabyte-scale data ingestion and versioning; Set technical direction for high-throughput data delivery to GPU/TPU fleets; Make system-level format and library design calls; Drive technical direction and interfaces across partner teams; Mentor senior engineers and lead design reviews; Shape platform roadmap for foundation models and multimodal workloads.
Seniority
Staff, hands-on IC with strategic influence
