Senior AI Engineer
Core
Build and operationalize machine learning models to detect anomalies, perform root cause analysis, and forecast hardware failures in network time-series data.
Role type
Senior IC machine learning engineer (network operations)
Builds
Production ML systems for network telemetry, predictive maintenance, and traffic forecasting
Domain
Telecommunications / Network Infrastructure
Deliverable
production ML models
Required skills
classical machine learning, time-series analysis, anomaly detection, model evaluation, production deployment, feature engineering, data pipeline development, network telemetry familiarity, MLOps lifecycle tooling
Preferred skills
Cisco platforms familiarity, streaming telemetry (Kafka, Pub/Sub), SDN experience, network operations experience
Technologies
BigQuery, Cloud Storage, Vertex AI, Kafka, Pub/Sub, OpenTelemetry, MLflow, Weights & Biases, PyTorch, SQL, Kubernetes
Responsibilities
Build and operationalize machine learning models to detect anomalies in network time-series data; Develop root cause analysis models to identify contributing factors and failure chains in network events; Create predictive maintenance models to forecast hardware failures and network degradation; Design evaluation frameworks that account for precision and recall, false-positive costs, and operational trust; Assess, clean, and structure network telemetry in GCP BigQuery, and build pipelines that convert it into machine-learning-ready features; Manage the model lifecycle, including experiment tracking, versioning, retraining pipelines, and production drift monitoring; Define model retraining triggers and health thresholds suited to network operations; Define the AI WAN closed-loop architecture, including model inputs, autonomous decisions, and actions requiring human approval; Build a human-in-the-loop recommendation layer before advancing to autonomous actions, and establish suitable guardrails, rollback mechanisms, and confidence thresholds
Seniority
Senior, hands-on IC
