Site Reliability Engineering Manager
Core
Lead site reliability engineering practices, infrastructure architecture, and incident response to ensure scalable and reliable services.
Role type
Senior IC Site Reliability Engineering Manager
Builds
Scalable, reliable infrastructure and services
Domain
Cloud Infrastructure / Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Infrastructure management, Capacity planning, Incident response, Root cause analysis, Automation, Scripting, Data analysis, Technical leadership, Budget management, Team coaching
Preferred skills
Leadership experience, Financial management experience, Advanced degree
Technologies
Oracle Security
Responsibilities
Design and architect infrastructure and services for reliability, guide capacity planning and resource forecasting, coordinate with development teams to build scalable infrastructure, coach team members on prototyping and testing, advise on data collection and technical analysis, support service monitoring and performance awareness, assist with incident response and root cause analysis, review health reports and recommend actions, oversee improvements to optimize deployments, build team knowledge of site reliability practices, create and own execution plans for projects, strengthen collaboration across teams, lead problem solving for operational issues, coach and develop team members.
Seniority
Senior, hands-on IC with leadership responsibilities
