NOC Engineer / SRE
Core
24/7 service reliability, incident response, operational automation, and observability at the intersection of traditional NOC and engineering-driven reliability practices.
Role type
Senior Individual Contributor SRE/NOC Engineer
Builds
Self-healing systems, automated remediation scripts, and improved incident response workflows
Domain
Cloud infrastructure, network operations, and reliability engineering
Required skills
Linux systems administration, incident management, cloud infrastructure (AWS), containers and orchestration (Docker, Kubernetes), scripting (Python, Bash, Go), networking fundamentals (DNS, TCP/IP, load balancing)
Preferred skills
SLO/SLI definition, Infrastructure as Code (Terraform, Ansible), security/compliance experience
Technologies
Grafana, Prometheus, Datadog, Splunk, CloudWatch, AWS, Azure, GCP, Kubernetes, Docker, Python, Bash, Go, Terraform, Ansible
Responsibilities
Act as primary or escalation responder in 24x7 on-call rotation, lead Major Incident response including triage and mitigation, design and maintain alerting strategies aligned with SLIs/SLOs, automate repetitive operational tasks to reduce toil, troubleshoot Linux-based systems and cloud platforms, support capacity planning and production release readiness
Seniority
Senior, hands-on IC