CareerPlanSign in

AI算力调度专家

Shenzhen, China💼 Full-time🗓 2026-08-07 → 2026-09-28

Core

Design and implement large-scale distributed scheduling systems for unified management and topology-aware scheduling of multi-generational, heterogeneous compute resources (NPU/CPU).

Role type

Senior IC distributed systems engineer (AI compute scheduling)

Builds

Distributed scheduling systems for training, inference, and RL sandbox scenarios

Domain

High-performance computing + AI infrastructure

Deliverable

production ML models

Required skills

Distributed systems architecture, multi-objective optimization algorithms, heterogeneous compute management, Kubernetes/Slurm custom scheduler development, RDMA network tuning, container image orchestration, high-concurrency system development

Preferred skills

Top-tier conference publications (OSDI/SOSP/EuroSys), reinforcement learning scheduling, complex workflow orchestration (Airflow/Argo/Ray), open-source community contributions

Technologies

Kubernetes, Slurm, Go, C++, Python, RDMA, Prometheus, Grafana, Ray, Airflow, Argo Workflows

Responsibilities

Design topology-aware scheduling for training checkpoints and data locality; develop mirror pre-warming and distributed caching strategies for inference; design task coordination mechanisms for RL sandbox environments; implement multi-tenant resource management with quota and preemption; optimize scheduler throughput and latency for 10k+ GPU clusters

Seniority

Senior, hands-on IC

Sourced via huawei · Listed on CareerPlan, which tracks 852,000+ jobs from 20+ sources.