Engineering Manager, ML Training Infrastructure
Core
Lead a team building and operating large-scale ML training infrastructure (GPU clusters, Kubernetes, orchestration) to support autonomous driving and mobility projects.
Role type
Senior Engineering Manager, ML Training Infrastructure
Builds
Scalable GPU compute environments, Kubernetes-based clusters, and workflow orchestration platforms for ML workloads
Domain
Automotive mobility / Autonomous driving / Cloud infrastructure
Deliverable
infrastructure
Required skills
Kubernetes orchestration, distributed systems, GPU cluster management, cloud platform architecture, ML workload optimization, system resilience, team leadership, technical roadmap planning
Preferred skills
Golang, automotive industry experience, security-focused environments
Technologies
Kubernetes, Ray, Airflow, Temporal, GPU fleets
Responsibilities
Lead a team of 7–8 engineers; define technical vision and roadmap for ML infrastructure; architect and operate resilient systems for massive ML training; manage GPU fleets and cluster scheduling; optimize workflow orchestration; partner with AD/ADAS and Woven City teams; enforce software practices and incident management.
Seniority
Senior, hands-on IC with leadership