CareerPlanSign in

Infra平台开发工程师-Data AML

北京💼 Full-time🗓 2026-09-28

Core

Building and managing AI infrastructure platforms for large-scale GPU clusters, distributed storage, and high-performance networks to support training and inference workloads.

Role type

Senior IC infrastructure engineer (AI platform)

Builds

Resource management, scheduling, and observability systems for AI training and inference

Domain

AI Infrastructure / Distributed Systems

Deliverable

infrastructure

Required skills

Go, Python, Java, distributed systems, microservices, observability, OLAP databases

Preferred skills

AIOps algorithms, large model/Agent development, monitoring and log platform construction

Technologies

Flask, Celery, Django, Tornado, NumPy, Gin, Gorm, Sarama, gRPC, OpenTSDB, Prometheus, Grafana, ELK, OpenTelemetry, ClickHouse

Responsibilities

Design resource management and scheduling for GPU clusters; Optimize resource efficiency and costs; Establish stability systems and incident response; Standardize development toolchains and CI/CD; Build intelligent AIOps capabilities for anomaly detection and automated remediation.

Sourced via bytedance · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.