Software Site Reliability Engineer - AI & Data Platforms
Core
Design, develop, and automate tools and frameworks to ensure reliability, scalability, and efficiency for large-scale distributed data platforms supporting analytics, reporting, and AI/ML applications.
Role type
Senior IC Site Reliability Engineer (Data Platforms)
Builds
Resilient data pipelines, monitoring systems, and automation tools for distributed data infrastructure
Domain
Enterprise technology, Big Data, Cloud Infrastructure, AI/ML infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, Golang, Java, distributed systems, data pipelines, cloud platforms (AWS/Azure/GCP), monitoring and alerting, root cause analysis, Kubernetes, Spark, Kafka, Flink, Airflow, DBT, data modeling, data warehousing, GPUs, MLFlow, LLMs
Preferred skills
Open source contributions, multi-tenant Kubernetes management, workflow orchestration, data structures & algorithms, software engineering best practices, secure coding, reusable frameworks
Technologies
Python, Golang, Java, AWS, Azure, Google Cloud Platform, Kubernetes, Spark, Kafka, Flink, Airflow, DBT, MLFlow, GPUs
Responsibilities
Design and build tools to improve reliability and scalability of distributed data systems; Implement advanced monitoring and alerting for on-prem and cloud workloads; Troubleshoot and resolve complex production incidents for critical applications; Collaborate with development teams to integrate reliability best practices; Proactively recommend improvements in architecture and operations
Seniority
Senior, hands-on IC