Software Engineering Manager, IS&T Ai & Data Platforms
Core
Lead a Site Reliability Engineering (SRE) team to ensure the reliability, scalability, and operational excellence of Apple's distributed data pipelines, analytics platforms, and enterprise data warehouse solutions.
Role type
Senior IC SRE Manager (hands-on technical leadership)
Builds
Scalable, resilient distributed data platforms and analytics solutions supporting real-time, near real-time, and batch processing.
Domain
Enterprise Data & Analytics / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering (SRE), cloud-native services, distributed systems, incident response, root cause analysis, Infrastructure as Code (IaC), cost optimization, team leadership, system design
Preferred skills
ETL frameworks (Apache Spark, Flink), messaging systems (Kafka), observability tools, modern distributed databases, data visualization tools, Generative AI for operations
Technologies
Kafka, Spark, Iceberg, Airflow, AWS, GCP, Kubernetes, Prometheus, Grafana, CloudWatch, Python, Java, Scala, Snowflake, Cassandra, SingleStore, SAP HANA, Tableau, Business Objects, ThoughtSpot
Responsibilities
Provide technical leadership and mentor an SRE team while actively contributing to code and platform reliability; Drive automation, Infrastructure as Code, and cost optimization initiatives; Participate in on-call rotations and lead incident response and post-mortem analysis; Collaborate with cross-functional teams to troubleshoot and enhance system reliability; Establish production readiness standards for security, disaster recovery, and documentation.
Seniority
Senior, hands-on IC with people management