Cloud Systems Engineer - Site Reliability
Core
Improve reliability, operability, and resilience of 24×7 SaaS production services and shared platforms supporting behavioral health practice management.
Role type
Senior Site Reliability Engineer (SRE)
Builds
High-availability, high-throughput, data- and compute-intensive critical systems for a 24×7 SaaS environment
Domain
Healthcare technology / Cloud Infrastructure
Required skills
Linux systems administration, cloud infrastructure design, containerization (Kubernetes), observability platform expertise, scripting (Bash/PowerShell/Python), infrastructure as code, incident response, distributed systems troubleshooting
Preferred skills
Azure cloud platform, Datadog observability, Prometheus/Grafana/New Relic, Terraform/OpenTofu, Ansible, prior software development experience
Technologies
Azure, Kubernetes, Datadog, Prometheus, Grafana, New Relic, Bash, PowerShell, Python, Terraform, OpenTofu, Ansible
Responsibilities
Design and maintain high-availability critical systems; define and improve reliability via SLIs/SLOs/error budgets; lead incident response and root cause analysis; automate operational toil; partner with development teams on system supportability; provide escalated technical guidance; ensure compliance with security and HIPAA policies.
Seniority
Senior, hands-on IC