Service Engineer
Core
Lead Azure's Incident Management practice, acting as the single point of command during high-severity incidents to restore services and protect customer trust.
Role type
Senior Incident Commander / Service Engineer (Cloud Operations)
Builds
Azure cloud infrastructure services and incident response frameworks
Domain
Cloud Operations / Site Reliability Engineering (SRE)
Required skills
Incident response and crisis management, cloud architecture patterns, microservices, containerization, monitoring and observability, automation scripting, ITIL frameworks, high availability and disaster recovery, root cause analysis, cross-functional leadership
Preferred skills
AI/ML integration in cloud infrastructure, chaos engineering, Windows/Linux debugging, cloud certifications (AWS/Azure/GCP), ITIL/SRE certifications
Technologies
Azure, AWS, GCP, Grafana, Prometheus, Datadog, Splunk, New Relic, PowerShell, Python
Responsibilities
Lead and manage high-severity incidents across Azure services as the central authority; drive incident reviews (RCAs/PIRs) and implement preventative improvements; collaborate with engineering and product teams to design resilient architecture; participate in on-call rotation; analyze telemetry to identify root causes; advocate for customer self-service capabilities and operational frameworks.
Seniority
Senior, hands-on IC with strategic oversight