机器学习平台研发工程师-Data
Core
Design and maintain GPU cluster scheduling systems to optimize resource utilization, reduce fragmentation, and ensure stable, scalable service for batch training, streaming updates, and online inference.
Role type
Senior IC machine-learning platform engineer (GPU scheduling)
Builds
High-availability GPU cluster scheduling infrastructure supporting batch and streaming ML workloads
Domain
Cloud infrastructure + Machine Learning
Deliverable
infrastructure
Required skills
GPU architecture knowledge, K8s, Docker, Ray, Apache YARN, Golang, C++, Python, dynamic scaling, fault tolerance design
Preferred skills
Thousand-card cluster scheduling experience
Responsibilities
Design multi-dimensional scheduling strategies to improve GPU utilization; Support complex scheduling needs for batch training, streaming updates, and online inference; Design high-availability fault tolerance and dynamic scaling mechanisms