Site Reliability Engineer - AI & Data Platforms
Core
Design, develop, and operate large-scale distributed data platforms supporting analytics, reporting, and AI/ML applications for Apple's enterprise and customer systems.
Role type
Senior Site Reliability Engineer (Data Platforms)
Builds
Reliable, scalable, and efficient distributed data platform systems and automation tools
Domain
Enterprise technology, Big Data, Cloud Infrastructure, AI/ML Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, Golang, Java, distributed systems, data pipelines, cloud platforms (AWS/Azure/GCP), monitoring and alerting, root cause analysis, Kubernetes, Spark, Kafka, Flink, Airflow, DBT, data modeling, data warehousing, GPUs, MLFlow, LLMs
Preferred skills
Open source contributions, multi-tenant Kubernetes management, workflow orchestration, data structures & algorithms, software engineering best practices, secure coding, reusable frameworks
Technologies
Python, Golang, Java, AWS, Azure, Google Cloud Platform, Kubernetes, Spark, Kafka, Flink, Airflow, DBT, MLFlow, GPUs
Responsibilities
Design and automate tools for reliability and scalability; implement advanced monitoring and alerting for on-prem and cloud workloads; troubleshoot and resolve complex production incidents for critical applications; collaborate with development teams to integrate reliability best practices; proactively recommend architectural and operational improvements
Seniority
Senior, hands-on IC
