Manager, Site Reliability Engineering
Core
Lead a team of SRE Automation Engineers to reduce manual toil and ensure high availability for a global Brokerage-as-a-Service platform.
Role type
Manager, Site Reliability Engineering (Hands-on IC + People Leadership)
Builds
Internal SRE platforms, automation workflows, and Kubernetes-based ecosystems for global financial markets
Domain
Fintech / Regulated Brokerage / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Team leadership & mentorship, Google SRE principles (SLOs, error budgets), Linux & TCP/IP networking, Kubernetes production management, AWS core services, Python or Golang, Rundeck and Airflow orchestration, Terraform & GitOps (ArgoCD), Incident response & RCA, Observability (Prometheus/Grafana), Security mindset (secrets management, supply chain)
Preferred skills
FinTech background with FIX/API connectivity, AI & Prompt Engineering (LLMs, Bedrock), Kafka/MQ/SQS middleware management
Technologies
Kubernetes, AWS, Terraform, ArgoCD, Rundeck, Airflow, Python, Golang, Prometheus, Grafana, Kafka, MQ, SQS
Responsibilities
Manage, mentor, and grow a team of SRE Automation Engineers; Lead design and development of automation tooling using Rundeck and Airflow; Define SLIs, SLOs, and error budgets; Set architectural standards for Infrastructure as Code; Review software architecture and Kubernetes metrics; Lead incident response and root-cause analysis
Seniority
Manager, hands-on IC with team leadership