Site Reliability Engineer, Enterprise Technology Services
Core
Design, build, and operate large-scale Identity Management Platform services ensuring high availability, security, and reliability for critical authentication, authorization, and provisioning across Apple's ecosystem.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Distributed identity management systems, automation frameworks, observability stacks, and resilience platforms.
Domain
Cloud infrastructure, Identity & Access Management (IAM), Distributed Systems
Required skills
Distributed systems architecture, SRE principles (SLIs/SLOs/SLAs), Incident management, Automation & CI/CD, Observability stack design, Security & compliance, Python/Java/Go programming, Capacity planning, Disaster recovery
Preferred skills
Chaos engineering, GenAI for alert engineering, Multi-cloud strategy, Event-driven architectures (Kafka), Cryptography & OAuth/SAML, Cyber security certifications
Technologies
Kubernetes, Helm, CRD, Prometheus, Grafana, OpenTelemetry, ELK, Datadog, Splunk, Kafka, RabbitMQ, Git
Responsibilities
Define and implement SLIs/SLOs/SLAs for platform reliability; Design and build resilient distributed systems with auto-scaling and failover; Lead incident response and post-mortems; Develop large-scale automation and tooling; Ensure security compliance and fraud prevention; Partner with engineering teams on system debugging and optimization.
Seniority
Senior, hands-on IC
