Site Reliability Engineer
Core
Own the operational health of Gamma's full backend platform, building automation and tooling to ensure reliability and observability for millions of daily users.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
Production backend systems, observability infrastructure, and automation tooling on AWS
Domain
SaaS, Cloud Infrastructure (AWS)
Required skills
AWS expertise, Python/Go/TypeScript, Infrastructure-as-Code (Terraform/CloudFormation), Observability (metrics/logging/tracing), Incident management, Distributed systems, Containerization (Docker/Kubernetes), Database optimization
Preferred skills
Kafka, Chaos engineering, Service mesh, Security/compliance frameworks (SOC 2/ISO 27001), AWS certifications
Technologies
AWS, Python, Go, TypeScript, Node.js, Terraform, CloudFormation, Docker, Kubernetes, Kafka
Responsibilities
Own reliability, availability, and performance of production systems; Build observability infrastructure; Design and ship automation to reduce toil; Lead incident response and post-mortems; Partner on architecture reviews and SLO/SLI design; Manage and optimize compute, networking, databases, and managed services
Seniority
Senior, hands-on IC