Site Reliability Engineer — Data Platforms, IS&T Ai & Data Platforms
Core
Operate and harden large-scale distributed data and ML platforms to ensure reliability, performance, and cost-efficiency for critical applications like analytics, reporting, and AI/ML.
Role type
Senior Site Reliability Engineer (Data Platforms)
Builds
Data processing ecosystems, distributed computing frameworks, and MPP query engines powering analytics and AI/ML workloads.
Domain
Cloud infrastructure, Big Data, Machine Learning
Deliverable
infrastructure
Required skills
Python, Go, Java, Scala, Bash, Kubernetes, Helm, Terraform, Pulumi, Spark, Flink, Trino, StarRocks, Unix/Linux, Incident Response, SLO/SLI management, Capacity Planning, CI/CD
Preferred skills
Multi-cloud operations, Open source contributions, Kafka Streams, Iceberg, Airflow, dbt, Multi-tenant Kubernetes management
Technologies
Kubernetes, Helm, Terraform, Pulumi, GitHub Actions, Jenkins, Spark, Flink, Trino, StarRocks, Kafka Streams, Iceberg, Airflow, dbt
Responsibilities
Drive down MTTR and automate operational toil, tune performance and cost, run capacity planning, root-cause production incidents, operate and harden big data platforms.
Seniority
Senior, hands-on IC
