CareerPlanSign in

大模型训练调度专家 - Seed Model

杭州💼 Full-time🗓 2026-09-28

Core

Design and develop machine learning system resource scheduling to support model training, evaluation, and inference across NLP, CV, and Speech scenarios.

Role type

Senior IC machine learning systems engineer (resource scheduling)

Builds

Distributed ML training clusters and inference services for large language models and multimodal applications

Domain

Artificial Intelligence / Distributed Systems / Cloud Infrastructure

Deliverable

infrastructure

Required skills

Linux environment development, Go/Python/Shell programming, Kubernetes architecture, container technologies (Docker/Containerd/Kata/Podman), distributed system principles, large-scale distributed system design, logical analysis and abstraction

Preferred skills

PyTorch/Megatron-LM frameworks, Ray framework, reinforcement learning frameworks, data-driven ML systems, large-scale AI task fault tolerance, high-performance computing, RDMA networks, storage systems, OS kernel, GPU hardware drivers, publications at OSDI/SOSP/NSDI/ATC/EuroSys

Technologies

Kubernetes, Docker, Containerd, Kata, Podman, PyTorch, Megatron-LM, Ray, RDMA

Responsibilities

Design and develop resource scheduling systems for ML training, evaluation, and inference; orchestrate heterogeneous resources (GPU, CPU, etc.) for stable and efficient usage; manage compute, RDMA network, and storage resources in large-scale clusters; handle scheduling across multi-datacenter, multi-region, and multi-cloud environments; optimize resource utilization through task prioritization, preemption, and queue management

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 862,000+ jobs from 20+ sources.