Site Reliability Engineer, Enterprise Technology Services
Core
Design, build, and scale a modern DevOps and SRE ecosystem from the ground up, establishing GitOps-driven cloud-native CI/CD platforms and reliability practices.
Role type
Senior Site Reliability Engineer (Platform Engineering)
Builds
Scalable, highly available infrastructure and cloud-native CI/CD platforms
Domain
Cloud infrastructure, DevOps, and Security Operations
Deliverable
infrastructure
Required skills
Java/JEE, Python, Bash, Lua, Terraform, Ansible, PL/SQL, Kubernetes, GitOps, AIOps, SecOps, TLS/mTLS
Preferred skills
AWS/GCP/Azure, Prometheus, Splunk, Grafana, CloudWatch, Machine Learning algorithms, Cryptography, Regular expressions, Vulnerability research
Technologies
Java, JEE, Python, Bash, Lua, Terraform, Ansible, PL/SQL, AWS, GCP, Azure, ScaleIO, Amazon S3, Prometheus, Splunk, Grafana, CloudWatch
Responsibilities
Engineer scalable, highly available infrastructure aligned with business and reliability objectives; Drive KPI, SLA, and SLO performance through measurable reliability and operational targets; Optimize system performance, capacity, cost, and resource utilization across growing workloads; Automate infrastructure, deployments, remediation, testing, and operational workflows using IaC and AI-driven automation; Expand observability through metrics, logs, traces, dashboards, alerting, and actionable service health insights; Apply AIOps capabilities to detect anomalies, correlate events, predict failures, and accelerate incident resolution; Harden systems through SecOps practices, vulnerability management, access controls, and continuous security improvements; Lead incident response, root-cause analysis, reliability improvements, and operational readiness for service expansion.
Seniority
Senior, hands-on IC with technical leadership