Senior Site Reliability Engineer - AI Platform
Core
Design and build foundational platform architectures from scratch to support scalable AI adoption and hyper-automation, establishing secure, scalable systems for teams to innovate independently.
Role type
Senior Site Reliability Engineer (AI Platform)
Builds
Foundational platform architectures, reliable infrastructure, CI/CD pipelines, and distributed systems for AI-powered automations.
Domain
AI Platform Engineering / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
SRE, software engineering, distributed systems, Infrastructure as Code, Kubernetes, containers, CI/CD, database modeling, applied AI (model behavior, costs, tokens), observability, security, performance optimization
Preferred skills
No-code/low-code platforms (Zapier, Base44, Botpress, Botcity), mentoring technical teams
Technologies
Kubernetes, containers, Vault, API Ingress, CI/CD tools, database systems
Responsibilities
Design and build foundational platform architectures from scratch; Develop reliable, secure, and highly scalable infrastructure; Establish platform guardrails and engineering standards; Build and maintain infrastructure using Infrastructure as Code; Develop and optimize CI/CD processes; Address high-scale performance challenges in distributed systems; Contribute to database architecture evolution; Apply software engineering best practices; Explore and incorporate applied AI capabilities; Collaborate within agile environments; Shape technical direction and mentor other engineers
Seniority
Senior, hands-on IC with technical leadership