Staff Site Reliability Engineer
Core
Design, build, and operate secure, resilient, scalable cloud infrastructure and platforms for large-scale customer-facing technology products.
Role type
Staff Site Reliability Engineer (Infrastructure)
Builds
Core infrastructure products, subsystems, and developer tooling
Domain
Cloud infrastructure, distributed systems, reliability engineering
Required skills
Cloud architecture, distributed systems, Terraform, Kubernetes/EKS, Go/Python, Redis/ElastiCache, observability, incident management, performance tuning, security mindset, code reviews, technical documentation
Preferred skills
AWS experience, mentoring engineers, influencing engineering practices
Technologies
AWS, Terraform, Kubernetes, EKS, Go, Python, Redis, ElastiCache, Prometheus, Grafana, OpenTelemetry
Responsibilities
Take end-to-end ownership of core infrastructure systems from design to production operations; define project goals and success metrics; translate requirements into practical designs; build secure, reliable, high-performing infrastructure; develop production software and developer tools; manage infrastructure via code and configuration; partner with product engineering to design for scale; participate in incident response and debugging; develop monitoring and observability practices; apply security-focused mindset; mentor engineers and drive collaboration across teams
Seniority
Staff, hands-on IC with mentorship and strategy influence
