CareerPlanSign in

云原生算力平台SRE工程师(深圳/北京/上海)

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Operate, troubleshoot, and optimize GPU/CPU heterogeneous computing infrastructure to ensure stable, efficient, and continuous compute output.

Role type

Senior Site Reliability Engineer (Compute Platform)

Builds

High-availability GPU/CPU compute clusters and automated operations tooling

Domain

Cloud-native computing infrastructure and heterogeneous hardware

Deliverable

production ML models | infrastructure

Required skills

GPU hardware and driver tuning, Kubernetes cluster management, Linux administration, Golang/Python/Java, heterogeneous computing (Mellanox/NCCL/Cuda)

Preferred skills

Experience with Docker, cloud-native disaster recovery design, automation and intelligent operations methods

Technologies

Kubernetes, Docker, Golang, Python, Java, Linux, CUDA, NCCL, Mellanox

Responsibilities

Daily operations and troubleshooting of GPU/CPU heterogeneous devices; Managing and governing K8s clusters including disaster recovery and security drills; Automating operational workflows for resource and change management.

Sourced via tencent · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.