Site Reliability Engineer
Core
Ensure reliability, scalability, and security of cloud platforms and SaaS systems operating at multi-region scale.
Role type
Site Reliability Engineer (SRE)
Builds
Cloud and data infrastructure, automation tools, and core data platform capabilities
Domain
Cloud infrastructure and SaaS platforms
Required skills
Kubernetes (EKS or self-managed), Docker, cloud platforms (AWS), programming/scripting (Python, Bash, Go), Infrastructure as Code (Ansible, Terraform), Linux systems, CI/CD pipelines, observability tools (ELK, Splunk, Prometheus)
Preferred skills
Cloud-based data platform management, AI-assisted SRE tooling, automation-first approaches
Technologies
Kubernetes, Docker, AWS, Python, Bash, Go, Ansible, Terraform, ELK, Splunk, Prometheus
Responsibilities
Build, deploy, and optimize cloud infrastructure; integrate software and systems engineering; monitor production systems and troubleshoot incidents; contribute to root cause analysis and postmortem reviews; collaborate with cross-functional teams to enhance operational efficiency through automation
Seniority
Mid-level, hands-on IC