CareerPlanSign in

机器学习平台研发工程师-Data

杭州💼 Full-time🗓 2026-09-28

Core

Design and maintain GPU cluster scheduling systems to optimize resource utilization, reduce fragmentation, and ensure stable, scalable service for batch training, streaming updates, and online inference.

Role type

Senior IC machine-learning platform engineer (GPU scheduling)

Builds

High-availability GPU cluster scheduling infrastructure supporting batch and streaming ML workloads

Domain

Cloud infrastructure + Machine Learning

Deliverable

infrastructure

Required skills

GPU architecture knowledge, K8s, Docker, Ray, Apache YARN, Golang, C++, Python, dynamic scaling, fault tolerance design

Preferred skills

Thousand-card cluster scheduling experience

Responsibilities

Design multi-dimensional scheduling strategies to improve GPU utilization; Support complex scheduling needs for batch training, streaming updates, and online inference; Design high-availability fault tolerance and dynamic scaling mechanisms

Sourced via bytedance · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.