AI算力调度专家
Core
Design and implement large-scale distributed scheduling systems for unified management and topology-aware scheduling of multi-generational, heterogeneous compute resources (NPU/CPU).
Role type
Senior IC distributed systems engineer (AI compute scheduling)
Builds
Distributed scheduling systems for training, inference, and RL sandbox scenarios
Domain
High-performance computing + AI infrastructure
Deliverable
production ML models
Required skills
Distributed systems architecture, multi-objective optimization algorithms, heterogeneous compute management, Kubernetes/Slurm custom scheduler development, RDMA network tuning, container image orchestration, high-concurrency system development
Preferred skills
Top-tier conference publications (OSDI/SOSP/EuroSys), reinforcement learning scheduling, complex workflow orchestration (Airflow/Argo/Ray), open-source community contributions
Technologies
Kubernetes, Slurm, Go, C++, Python, RDMA, Prometheus, Grafana, Ray, Airflow, Argo Workflows
Responsibilities
Design topology-aware scheduling for training checkpoints and data locality; develop mirror pre-warming and distributed caching strategies for inference; design task coordination mechanisms for RL sandbox environments; implement multi-tenant resource management with quota and preemption; optimize scheduler throughput and latency for 10k+ GPU clusters
Seniority
Senior, hands-on IC