系统智能运维专家
Core
Design and implement stability risk interception, fault diagnosis, and self-healing systems for million-scale data centers; lead long-term technical planning for intelligent O&M systems.
Role type
Senior IC system architecture engineer (stability & intelligent O&M)
Builds
High-availability, high-reliability, and secure intelligent O&M systems for large-scale data centers
Domain
Cloud computing, distributed systems, and data center operations
Deliverable
production ML models
Required skills
System resilience and high availability design, Cloud computing, Containerization, Microservices architecture, Linux OS, Large-scale distributed system design, Fault diagnosis and troubleshooting
Preferred skills
SRE operational scenarios, Intelligent O&M system architecture
Technologies
Kubernetes, Docker, Prometheus, Grafana, ELK Stack
Responsibilities
Design fault diagnosis and self-healing systems for million-scale data centers; Collaborate with SRE and business teams to solve production pain points; Lead long-term technical planning for O&M systems; Participate in emergency response for complex production faults
Seniority
Senior, hands-on IC