Site Reliability Engineer — ETL Platform, IS&T Ai & Data Platforms
Core
Own reliability, performance, data freshness, and operational readiness for a production ETL platform supporting data ingestion, transformation, and loading workflows.
Role type
Senior Site Reliability Engineer (Data Infrastructure)
Builds
Production ETL pipelines, Kubernetes-based loader jobs, Spark-on-EKS workloads, and Datalake/Lakehouse data loads.
Domain
Data Infrastructure / Cloud Platforms
Deliverable
production ML models
Required skills
Kubernetes/EKS operations, Apache Spark tuning, Apache Airflow, Kafka streaming, Linux troubleshooting, SQL, Python scripting, Bash scripting, observability tooling, GitOps, incident management, root cause analysis.
Preferred skills
Metadata-driven ETL platforms, Lakehouse technologies (Iceberg, Hive Metastore), infrastructure-as-code (Terraform, Helm, Argo CD), multi-region data platforms, PKI/TLS management, Spark performance tuning at scale.
Responsibilities
Operate and triage production pipelines end-to-end (extractors, loaders, batch jobs, streaming ingestion, Spark workloads, Airflow DAGs), tune Kubernetes and Spark for stability, build observability tooling, drive root cause analysis, write automation to reduce toil, manage configuration and secrets through GitOps-style processes.
Seniority
Senior, hands-on IC