Lead Software Engineer, Cloud Site Reliability (SRE)
Core
Lead 24x7 cloud reliability operations and incident management for critical Azure-based technology environments.
Role type
Lead Software Engineer, Cloud Site Reliability (SRE)
Builds
Resilient, scalable cloud-native environments on Azure with high availability and SLA adherence
Domain
Cloud Infrastructure / Site Reliability Engineering
Required skills
Azure IaaS, AKS, Kubernetes, Docker, Datadog, Azure Monitor, Terraform, ARM templates, Helm, PowerShell, Python, Bash, ServiceNow, Incident Management, Root Cause Analysis, Operational Reporting
Preferred skills
Multi-cloud (AWS), AIOps, Predictive Monitoring, Anomaly Detection, Self-healing systems, Azure/Datadog/Kubernetes certifications
Technologies
Azure, AKS, Kubernetes, Docker, Datadog, Azure Monitor, Terraform, ARM templates, Helm, Power Automate, ServiceNow, Power BI
Responsibilities
Lead 24x7 NOC operations through rotational shifts; Act as Major Incident Manager for P1/P2 incidents; Manage and troubleshoot Azure infrastructure; Administer AKS, Kubernetes, and Docker environments; Establish observability practices using logs, metrics, and traces; Drive proactive monitoring and AIOps initiatives; Build automation and self-healing workflows; Collaborate on deployment pipelines and cloud-native architecture; Develop operational dashboards and reports; Lead monthly business reviews; Mentor team members and standardize processes
Seniority
Lead, hands-on IC with mentorship responsibilities

