Engineering Manager, Site Reliability Engineering
Core
Lead a team of Site Reliability Engineers to build and operate large-scale, fault-tolerant distributed systems for Google Cloud, balancing feature development with production stability.
Role type
Senior IC manager (Site Reliability Engineering)
Builds
Large-scale, massively distributed, fault-tolerant systems for Google Cloud
Domain
Cloud infrastructure, distributed systems, SRE
Required skills
Software development, distributed systems design, people management, project leadership, SRE principles, automation, incident command
Preferred skills
Managing high-performing teams for 24/7 systems, cross-functional project leadership, architectural decision influence
Technologies
None explicitly listed beyond general distributed systems
Responsibilities
Lead and grow a team of SREs, Partner with Software Development teams to design reliable services, Drive evolution of monitoring and incident response frameworks, Implement automation to eliminate manual tasks, Participate in on-call rotations and lead incident command