Infra平台开发工程师-Data AML
Core
Building and managing AI infrastructure platforms for large-scale GPU clusters, distributed storage, and high-performance networks to support training and inference workloads.
Role type
Senior IC infrastructure engineer (AI platform)
Builds
Resource management, scheduling, and observability systems for AI training and inference
Domain
AI Infrastructure / Distributed Systems
Deliverable
infrastructure
Required skills
Go, Python, Java, distributed systems, microservices, observability, OLAP databases
Preferred skills
AIOps algorithms, large model/Agent development, monitoring and log platform construction
Technologies
Flask, Celery, Django, Tornado, NumPy, Gin, Gorm, Sarama, gRPC, OpenTSDB, Prometheus, Grafana, ELK, OpenTelemetry, ClickHouse
Responsibilities
Design resource management and scheduling for GPU clusters; Optimize resource efficiency and costs; Establish stability systems and incident response; Standardize development toolchains and CI/CD; Build intelligent AIOps capabilities for anomaly detection and automated remediation.