Reliability Engineer 4 (Observability Specialist )
Core
Senior Reliability Engineer establishing observability governance, defining SLIs/SLOs, and ensuring production readiness for enterprise applications.
Role type
Senior IC Reliability Engineer (Observability)
Builds
Production-ready enterprise applications and services with robust observability and reliability metrics.
Domain
Financial Services / Observability Engineering
Required skills
SLI/SLO definition, Observability Governance, Telemetry Analysis, Incident Analysis, Technical Leadership, Distributed Systems Architecture
Preferred skills
APM/RUM/Synthetics expertise, Cloud/Kubernetes proficiency, Alert Governance, RCA
Technologies
Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, OpenTelemetry
Responsibilities
Lead definition and governance of SLIs, SLOs, and error budgets; Design scalable observability architectures and instrumentation standards; Develop executive and operational service health dashboards; Analyze telemetry and incident data to identify gaps and improve detection; Provide technical mentorship on monitoring design patterns and alert governance; Partner with SRE and product teams to ensure applications are fully instrumented.
Seniority
Senior, hands-on IC