Site Reliability Engineer
Core
Design, build, and scale a modern DevOps and SRE ecosystem from scratch, managing the lifecycle of machine learning models in production and non-production environments while ensuring scalability, reliability, and performance of Apple B2B systems.
Role type
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Builds
GitOps-driven, cloud-native CI/CD platforms, Kubernetes clusters, and ML model lifecycle pipelines
Domain
B2B supply chain integrations, cloud-native infrastructure, machine learning operations
Deliverable
production ML models
Required skills
Kubernetes (EKS/AKS/GKE/OpenShift), GitOps (ArgoCD/Flux), Infrastructure as Code, Python, Java, distributed systems, observability (Splunk/Grafana/Prometheus), database tuning (Oracle/MongoDB), security protocols
Preferred skills
Container orchestration internals, service mesh (Istio/Linkerd), multi-cloud/hybrid environments, SLO/SLI design, error budgets, middleware platforms
Technologies
Kubernetes, ArgoCD, Flux, Python, Java, Splunk, Grafana, Prometheus, Oracle, MongoDB, Istio, Linkerd
Responsibilities
Design and build DevOps platforms end-to-end, lead incident response and RCA, implement telemetry and monitoring solutions, perform performance tuning of applications and databases, manage on-call and incident management for internet-facing services
Seniority
Senior, hands-on IC with architectural leadership
