Staff Site Reliability Engineer
Core
Design, build, and maintain cloud infrastructure and observability systems to ensure the reliability, scalability, and performance of ServiceTitan's applications.
Role type
Staff Site Reliability Engineer (IC)
Builds
Kubernetes-based compute platform, observability dashboards, alerting systems, and automation tools for cloud infrastructure.
Domain
Cloud Infrastructure & Site Reliability Engineering
Required skills
Kubernetes, SRE principles (SLIs/SLOs), Cloud networking (AWS/Azure/GCP), Observability stacks, CI/CD, Distributed systems troubleshooting, AI-assisted engineering tools
Preferred skills
Experience with .NET, Python, or Java web applications, Experience with AI agents for root cause analysis
Technologies
Kubernetes, Azure, AWS, OpenTelemetry, Prometheus, Grafana, Datadog, Elasticsearch, GitHub Actions, Claude Code, GitHub Copilot
Responsibilities
Diagnose and resolve production incidents using runbooks, Design and maintain observability dashboards grounded in SLIs/SLOs, Operate and improve Kubernetes-based compute platform, Build automation to reduce manual operational work, Partner with product engineering teams on architecture reviews, Drive adoption of reliability best practices across teams
Seniority
Staff, hands-on IC with strategic impact
