Observability SRE Manager, Apple Services Engineering
Core
Senior leadership role owning the reliability, scalability, and performance of Apple's observability platform (metrics, logging, tracing, alerting) and driving the evolution of the SRE practice with a focus on AI integration.
Role type
Senior SRE Manager / Technical Leader
Builds
Observability platform infrastructure and SRE practices for Apple Services Engineering
Domain
Cloud Infrastructure / Observability / AI in Operations
Deliverable
production ML models
Required skills
Technical leadership, people management, observability systems (metrics, logging, tracing, alerting, SLOs), distributed systems, Kubernetes, shell scripting, Python, AI/ML tooling for operations, cross-functional initiative driving
Preferred skills
Prometheus ecosystem, cloud-native stacks (Thanos, Splunk, OpenTelemetry), cloud platforms (AWS, GCP, Azure), Infrastructure as Code (Terraform, Ansible), open-source orchestration (Helm, Puppet, Spinnaker), TCP/IP, web application security
Technologies
Kubernetes, Prometheus, Thanos, Splunk, OpenTelemetry, Terraform, Ansible, Helm, Puppet, Spinnaker, Python, Go, Java, Scala, AWS, GCP, Azure
Responsibilities
Define technical direction and strategic roadmap for observability infrastructure; lead and mentor the SRE team; drive automation and AI adoption to reduce toil; manage on-call rotations and incident response; partner with engineering teams to influence system design for reliability; communicate strategy and progress to executive leadership
Seniority
Senior, hands-on IC with management responsibilities