大模型训练调度专家 - Seed Model
Core
Design and develop machine learning system resource scheduling to support model training, evaluation, and inference across NLP, CV, and Speech scenarios.
Role type
Senior IC machine learning systems engineer (resource scheduling)
Builds
Distributed ML training clusters and inference services for large language models and multimodal applications
Domain
Artificial Intelligence / Distributed Systems / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Linux environment development, Go/Python/Shell programming, Kubernetes architecture, container technologies (Docker/Containerd/Kata/Podman), distributed system principles, large-scale distributed system design, logical analysis and abstraction
Preferred skills
PyTorch/Megatron-LM frameworks, Ray framework, reinforcement learning frameworks, data-driven ML systems, large-scale AI task fault tolerance, high-performance computing, RDMA networks, storage systems, OS kernel, GPU hardware drivers, publications at OSDI/SOSP/NSDI/ATC/EuroSys
Technologies
Kubernetes, Docker, Containerd, Kata, Podman, PyTorch, Megatron-LM, Ray, RDMA
Responsibilities
Design and develop resource scheduling systems for ML training, evaluation, and inference; orchestrate heterogeneous resources (GPU, CPU, etc.) for stable and efficient usage; manage compute, RDMA network, and storage resources in large-scale clusters; handle scheduling across multi-datacenter, multi-region, and multi-cloud environments; optimize resource utilization through task prioritization, preemption, and queue management
Seniority
Senior, hands-on IC