AIOps开发工程师-基础设施
Core
Building a comprehensive AIOps platform for network observability, intelligent fault diagnosis, and automated remediation across physical and virtual networks.
Role type
Senior IC AIOps Engineer (Network Infrastructure & ML)
Builds
Streaming telemetry data pipelines, intelligent root cause analysis systems, LLM-based运维 assistants, and capacity prediction models.
Domain
Telecommunications / Network Infrastructure / AIOps
Deliverable
production ML models
Required skills
Golang or Python, Linux network protocol stack, Spine-Leaf Fabric architecture, EVPN/VXLAN, BGP/OSPF, Microservices, Docker/Kubernetes, CI/CD, Kafka, Flink, ClickHouse/TSDB, Prometheus/OpenTelemetry, Neo4j, Machine Learning (anomaly detection, event correlation), LLM/Agent development (RAG, tool calling)
Preferred skills
Experience in real-time data pipeline construction, observability platform development, applying ML to network operations, LLM application in IT operations
Technologies
Golang, Python, Linux, Spine-Leaf, EVPN, VXLAN, BGP, OSPF, Kafka, Flink, ClickHouse, TSDB, Prometheus, OpenTelemetry, Neo4j, RAG, Kubernetes
Responsibilities
Design and develop high-availability data pipelines integrating multi-source telemetry (GNMI, NETCONF, IPFIX, SNMP); Implement ML algorithms for anomaly detection and root cause analysis across the full network stack; Develop and deploy LLM-based agents for automated troubleshooting and runbook execution; Build predictive models for network capacity and risk management; Ensure stability and performance of the end-to-end AIOps engineering system.