Site Reliability Engineer - Insights
Core
Build and manage large-scale, highly available Big Data ecosystems supporting exabytes of data for Apple's global manufacturing operations.
Role type
Senior Site Reliability Engineer (Data Infrastructure & AIOps)
Builds
Exabyte-scale data platforms, analytics tools, and AIOps capabilities for manufacturing services
Domain
Manufacturing operations, Big Data, Cloud Infrastructure
Deliverable
production ML models
Required skills
Python, Java, AWS/GCP, Kafka, Elasticsearch, Redis, Bigtable, AI/ML, LLM techniques, anomaly detection, predictive alerting
Preferred skills
Microservices, Docker, Kubernetes, ArgoCD, Jenkins, GitHub Actions, Grafana, Prometheus, Kibana, MySQL, PostgreSQL
Technologies
Kafka, Elasticsearch, Redis, Bigtable, Druid, ClickHouse, Docker, Kubernetes, ArgoCD, Jenkins, GitHub Actions, Grafana, Prometheus, Kibana, MySQL, PostgreSQL
Responsibilities
Own reliability, performance, and scalability of services within the Insight ecosystem; Drive automation to reduce operational toil; Instrument services for deep observability with dashboards, alerts, and SLOs; Partner with incident management for response and post-incident reviews; Build and advance AIOps capabilities including AI-driven alerting and automated triage
Seniority
Senior, hands-on IC