机器学习生产管理研发工程师 - Data AML
Core
Ensuring stability, governance, and efficiency of the Applied Machine Learning (AML) training platform for internal business units (Douyin, Toutiao) and external clients via Volcano Engine.
Role type
Senior IC Machine Learning Infrastructure Engineer (Training Platform)
Builds
Distributed training frameworks, task scheduling systems, and resource management tools for ML workloads.
Domain
Internet / Machine Learning Infrastructure / Cloud Computing
Deliverable
infrastructure
Required skills
Linux, Python, Go, Kubernetes, Distributed Systems, Root Cause Analysis, SLO/SLA definition, Resource Optimization
Preferred skills
TensorFlow/PyTorch inference principles, HDFS/BMQ, Cross-datacenter networking, Chaos Engineering, Profiling
Responsibilities
Manage end-to-end training task stability from scheduler to cluster execution; Lead on-call response and incident post-mortems; Optimize resource utilization (GPU/CPU/network) and implement throttling strategies; Implement automated pre-flight checks and change risk control mechanisms.
Seniority
Senior, hands-on IC