Senior Site Reliability Engineer
Core
Build and scale a modern DevOps and SRE ecosystem from scratch, managing the lifecycle of machine learning models in production and non-production environments while ensuring system reliability and performance.
Role type
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Builds
GitOps-driven, cloud-native CI/CD platforms, Kubernetes clusters, and ML model pipelines
Domain
B2B supply chain integrations, cloud infrastructure, and distributed systems
Deliverable
production ML models
Required skills
Kubernetes, GitOps (ArgoCD/Flux), Infrastructure as Code, Python, Java, observability (Splunk/Grafana/Prometheus), distributed systems, incident management, performance tuning
Preferred skills
NoSQL databases, relational databases, container orchestration internals, service mesh (Istio/Linkerd), multi-cloud environments, SLO/SLI design, error budgets
Technologies
Kubernetes, EKS, AKS, GKE, OpenShift, ArgoCD, Flux, Splunk, Grafana, Prometheus, Python, Java, Oracle, MongoDB, Cassandra, DynamoDB, Istio, Linkerd, WebMethods
Responsibilities
Design and build end-to-end DevOps platforms, operate Kubernetes platforms, lead incident response and RCA, implement GitOps-based deployment models, establish IaC practices, support internet-facing production services, tune application and database performance
Seniority
Senior, hands-on IC with architectural leadership
