Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Core
Operate, monitor, and triage production and non-production environments for Apple's multi-cloud data platform, supporting big data pipelines and ML/AI services.
Role type
Senior Site Reliability Engineer (Multi-Cloud Infrastructure)
Builds
Reliable multi-cloud infrastructure (AWS, GCP, on-prem Kubernetes) for internal data and AI product teams.
Domain
Cloud Infrastructure / Data Platform / AI/ML
Deliverable
infrastructure
Required skills
AWS core services (IAM, EKS, RDS, S3, VPC), Kubernetes administration, Python, Infrastructure-as-Code (Terraform, Crossplane), GitOps (Flux), Prometheus, Grafana, Splunk
Preferred skills
Golang, multi-cloud migration scenarios, Spark/Flink on Kubernetes, automation tooling
Technologies
AWS, GCP, Kubernetes, Terraform, Crossplane, Flux, Prometheus, Grafana, Splunk, Python, Golang
Responsibilities
Operate and monitor production/non-production environments across the ADP portfolio; Participate in rotating on-call schedules; Own operational health of multi-cloud infrastructure as SME; Provide Slack-based support to internal customers; Debug production incidents involving IAM, storage, and cluster disruptions; Partner with dev teams to design monitoring and alerting; Maintain and evolve IaC and GitOps workflows; Build automation and self-healing tooling.
Seniority
Mid-level to Senior, hands-on IC
