Staff Site Reliability Engineer
Core
Own end-to-end design, development, deployment, and operation of critical cloud infrastructure systems for large-scale customer-facing products.
Role type
Senior Staff Site Reliability Engineer (Infrastructure)
Builds
Secure, resilient, scalable, and cost-efficient cloud infrastructure and developer tooling
Domain
Cloud Infrastructure / Distributed Systems
Required skills
Cloud architecture, Terraform, Kubernetes/EKS, Go/Python, Redis/ElastiCache, Prometheus/Grafana/OpenTelemetry, incident response, performance tuning, security mindset, code reviews, technical documentation
Preferred skills
AWS experience, distributed systems expertise, observability practices
Technologies
Terraform, AWS, Kubernetes, EKS, Go, Python, Redis, ElastiCache, Prometheus, Grafana, OpenTelemetry
Responsibilities
Design and deploy production software and developer-facing tools; manage infrastructure through code and configuration; participate in incident response and debugging; develop monitoring and observability practices; mentor engineers and influence technical standards; drive collaboration across engineering and stakeholder groups
Seniority
Senior, hands-on IC with mentorship responsibilities