AIML - Senior ML/RL Training Infrastructure Engineer, AFM
Core
Design, build, and scale high-performance distributed training infrastructure for Apple's foundation models, specifically focusing on large-scale reinforcement learning (RL) workloads.
Role type
Senior IC ML/RL Training Infrastructure Engineer
Builds
Robust, high-performance RL pipelines and training frameworks for Apple's foundation models
Domain
Artificial Intelligence / Reinforcement Learning / High-Performance Computing
Deliverable
production ML models
Required skills
Distributed systems design, TPU-based training, JAX, PyTorch, cluster orchestration, performance profiling, fault tolerance, actor/learner architectures, experience replay, parallel environment execution
Preferred skills
Python software engineering, XLA internals, debugging GPU/TPU architectures, building training services, CI/CD practices, cloud-scale cluster experience, custom hardware expertise
Technologies
JAX, PyTorch, TPU, XLA, Python, GPU
Responsibilities
Design and optimize distributed RL training pipelines; tune low-level performance and compiler settings; manage cluster-level orchestration and resource allocation; ensure pipeline reliability and observability; diagnose bottlenecks in large-scale ML jobs; collaborate with researchers to accelerate experimentation
Seniority
Senior, hands-on IC