Senior Site Reliability Engineer
Core
Improve reliability, resilience, and operational effectiveness of cloud-based technology platforms and services by implementing automation, monitoring, and infrastructure improvements.
Role type
Senior Site Reliability Engineer (IC)
Builds
Automated CI/CD pipelines, Infrastructure as Code configurations, monitoring and observability systems, and resilient cloud infrastructure.
Domain
Cloud Infrastructure & Site Reliability Engineering
Required skills
Cloud platform operations (AWS/Azure), Infrastructure as Code (Terraform/CloudFormation), Python/Go/Ruby scripting, Linux/Unix systems administration, CI/CD pipeline design, Incident investigation and root cause analysis, Distributed systems knowledge, Observability tooling (Prometheus/Grafana/OpenTelemetry), Security practices in cloud and CI/CD, Mentorship
Preferred skills
Automated recovery and resilience engineering, Self-service platform capabilities, Large-scale distributed systems operations
Technologies
AWS, Microsoft Azure, Terraform, CloudFormation, Python, Ruby, Go, Linux/Unix, Prometheus, Grafana, OpenTelemetry, Containers
Responsibilities
Design and implement automation to reduce manual operational activities; Build and improve monitoring, logging, alerting, and observability; Investigate complex production issues and drive improvements; Develop and maintain CI/CD pipelines; Improve cloud infrastructure security and operational practices; Mentor engineers across teams
Seniority
Senior, hands-on IC