SRE-Azure - Global Industrial
Core
Improve reliability, availability, scalability, and performance of enterprise applications and platforms hosted in Microsoft Azure using software engineering, cloud infrastructure, and automation.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Resilient, cloud-native solutions on Microsoft Azure
Domain
Cloud Infrastructure / Site Reliability Engineering
Required skills
Azure services, Kubernetes (AKS), Infrastructure as Code (Terraform), CI/CD pipelines, observability tools, incident management, root cause analysis, capacity planning, distributed systems architecture
Preferred skills
None explicitly stated
Technologies
Microsoft Azure, Kubernetes, Terraform, Azure DevOps, GitHub Actions, Azure Monitor, Log Analytics, Application Insights, Grafana, Datadog, Dynatrace
Responsibilities
Monitor and optimize system performance and availability; Define and manage SLIs, SLOs, SLAs, and error budgets; Lead incident response and root cause analysis; Automate operational processes and deployments; Design and support Azure infrastructure including AKS and networking; Partner with development teams on release validation and production readiness
Seniority
Mid-to-Senior, hands-on IC