企业微信-机器学习平台调度工程师-(成都/北京)
Core
Design and optimize global resource scheduling for large-scale GPU clusters to ensure efficient and stable execution of offline and online distributed training tasks.
Role type
Senior IC machine-learning platform scheduling engineer
Builds
High-availability scheduling frameworks supporting distributed training on Kubernetes and cloud-native environments
Domain
Cloud-native infrastructure, high-performance computing, distributed machine learning
Deliverable
infrastructure
Required skills
Go, Python, C++, Kubernetes core components, RDMA, distributed storage, deep learning frameworks, OpenMP, MPI
Preferred skills
Hybrid cloud, virtualization, heterogeneous computing
Responsibilities
Lead global resource scheduling for large-scale GPU clusters; Optimize RDMA networks and distributed storage for training performance; Develop K8s schedulers, CSI plugins, and CRDs for task orchestration and disaster recovery; Explore hybrid cloud and virtualization technologies.
Seniority
Senior, hands-on IC