Site Reliability Engineer, Apple Data Platform / Big Data Platform
Core
Operate, monitor, and triage production and non-production environments for Apple's multi-cloud data platform, supporting internal teams building data and AI products.
Role type
Senior Site Reliability Engineer (Big Data Platform)
Builds
Multi-cloud infrastructure, big data engines (Spark, Flink, Airflow, Trino), ML/AI platform services, and data governance tooling.
Domain
Big Data, Cloud Infrastructure, Machine Learning Operations
Deliverable
production ML models | infrastructure
Required skills
Python, Kubernetes, AWS or GCP, Big Data technologies (Spark, Flink, Airflow, Trino), SRE principles, incident response, automation scripting
Preferred skills
Golang, REST Catalog services (Glue), data governance frameworks, observability tooling (Prometheus, Grafana, Splunk), CI/CD pipelines, S3/cloud storage fundamentals
Technologies
Spark, Flink, Airflow, Trino, Kubernetes, AWS, GCP, Prometheus, Grafana, Splunk, PagerDuty, Glue Catalog
Responsibilities
Operate and monitor production/non-production environments; participate in rotating on-call schedules; own operational health of big data services as SME; serve as primary point of contact for internal customers via Slack; screen and resolve support tickets; partner with dev teams to design monitoring and alerting; build automation and self-healing tooling; escalate and resolve production issues.
Seniority
Mid-level, hands-on IC
