Senior Site Reliability Engineer II
Core
Improving reliability, availability, performance, and operational quality of production systems across cloud and on-premises environments.
Role type
Senior Site Reliability Engineer (IC)
Builds
Cloud and on-premises infrastructure, Kubernetes clusters, automation scripts, monitoring dashboards, and disaster recovery procedures.
Domain
Cloud Infrastructure & Site Reliability Engineering
Required skills
Incident response and root-cause analysis, Kubernetes and containerized workloads, Linux/UNIX and Windows administration, Infrastructure as Code (IaC), Python and Shell scripting, observability and log analysis, automation and system provisioning, disaster recovery planning.
Preferred skills
Experience leading incident reviews, cross-functional collaboration with security and development teams, system modernization projects.
Technologies
Kubernetes, Linux, Windows, Python, Shell, PowerShell, IaC, monitoring tools, logging systems.
Responsibilities
Lead incident response, postmortems, and gap assessments; monitor environments and respond to alerts; design and maintain automation and runbooks; establish logging and alerting standards; build dashboards for system health; plan and implement changes and service requests; participate in disaster-recovery exercises.
Seniority
Senior, hands-on IC