Staff Site Reliability Expert
Core
Shape infrastructure reliability, developer experience, and platform capabilities for a global commerce platform serving merchants.
Role type
Staff Site Reliability Engineer (Platform)
Builds
Scalable, automated AWS infrastructure and platform services enabling product teams to ship independently.
Domain
Cloud Infrastructure / SaaS Commerce
Required skills
AWS operations, Infrastructure as Code (Terraform), Container orchestration (Docker, Kubernetes, ECS), Linux systems, Python/Ruby/Go programming, Shell scripting, Observability, Incident management, Disaster recovery, Cloud cost optimization, Relational/NoSQL databases (MySQL, PostgreSQL, Redis, DynamoDB)
Preferred skills
SaaS environment experience, Agile/Continuous delivery practices, Systems thinking, Mentorship
Technologies
AWS, Terraform, Docker, Kubernetes, ECS, Python, Ruby, Go, Shell, MySQL, PostgreSQL, Redis, DynamoDB
Responsibilities
Design and maintain scalable AWS infrastructure using IaC; Build tools and platform services for product engineers; Partner on resilient, secure, cost-effective system design; Improve engineering practices across distributed teams; Champion observability, HA, incident management, and DR; Lead incident response and drive lasting improvements; Mentor engineers and participate in on-call rotation.
Seniority
Staff, hands-on IC with strategic influence