Site Reliability Engineer, Enterprise Technology Services
Core
Design, build, and operate large-scale Identity Management Platform services ensuring high availability, reliability, and security for authentication, authorization, and user provisioning across Apple's ecosystem.
Role type
Senior Site Reliability Engineer (Identity Management)
Builds
Distributed identity services, automation frameworks, observability stacks, and resilience platforms for critical authentication and authorization transactions.
Domain
Identity Management, Cloud Infrastructure, Distributed Systems
Required skills
Java, Python, Go, Bash, Ansible, Kubernetes, Prometheus, Grafana, OpenTelemetry, Kafka, Relational/NoSQL databases, Chaos Engineering, Incident Management, Security Compliance (ISO-27001, PCI)
Preferred skills
Machine Learning for anomaly detection, Generative AI for alert engineering, Infrastructure as Code (Helm, CRD), OAuth/SAML/SSO, Cybersecurity certifications
Technologies
Java, Python, Go, Bash, Ansible, Kubernetes, Prometheus, Grafana, OpenTelemetry, ELK, Splunk, Kafka, RabbitMQ, Helm, CRD
Responsibilities
Define and implement SLIs, SLOs, and SLAs for platform reliability; Design and manage distributed systems with capacity planning and disaster recovery; Lead incident response and post-mortems; Develop large-scale automation and CI/CD pipelines; Ensure security posture and compliance with industry standards; Partner with engineering teams to optimize system performance and architecture.
Seniority
Senior, hands-on IC
