Site Reliability Engineer - Data Infrastructure
Core
Engineering resilience, scalability, and efficiency for core data services and underlying platforms powering products.
Role type
Site Reliability Engineer (Data Infrastructure)
Builds
Resilient data infrastructure, AI infrastructure, and data center environments
Domain
Data Infrastructure / Cloud Systems
Deliverable
infrastructure
Required skills
Linux, networking, scripting (Python, Bash, Go), incident response, change management, automation, observability
Preferred skills
Kubernetes, Docker, data stores (MySQL, Redis, PostgreSQL), monitoring tools (Prometheus, Grafana, ELK), data center operations
Technologies
Kubernetes, Redis, MySQL, Message Queue, Python, Go, Bash, Docker, Prometheus, Grafana, ELK Stack
Responsibilities
Respond to and resolve production incidents; perform deployments and configuration changes; automate repetitive tasks; refine monitoring dashboards and alert thresholds; support data center and AI infrastructure operations
Seniority
Mid-level, hands-on IC