Principal Site Reliability Engineer, Compute Infrastructure
Core
Build software engines, intelligent pipelines, and autonomous systems to power cloud presence and treat the cloud as a programmable, AI-orchestrated entity.
Role type
Principal Site Reliability Engineer (Compute Infrastructure)
Builds
Highly available cloud applications across AWS and GCP, autonomous agents for predictive auto-scaling and self-healing, and globally distributed Kubernetes fleets.
Domain
Cloud Infrastructure, Multi-cloud (AWS, GCP, Azure), AI/LLM integration
Deliverable
production ML models
Required skills
Cloud Software Engineering, Distributed Systems Infrastructure, TypeScript (Node.js), Go, Python, Kubernetes, Multi-cloud platforms
Preferred skills
LLM integration, Vector databases, Prompt engineering, AI-assisted code reviews, GitHub Actions, GitLab CI
Technologies
AWS, GCP, Azure, Kubernetes, Go, Python, TypeScript, LLM APIs, Vector databases
Responsibilities
Lead enterprise architectural strategy for multi-cloud integration of AI workflows; Architect telemetry pipelines using LLMs for IAM auditing and vulnerability patching; Engineer autonomous agents for predictive auto-scaling and self-healing; Act as authority for complex anomalies leading AI-driven root-cause analysis; Define technical vision for globally distributed Kubernetes fleets with AI-driven traffic routing.
Seniority
Principal, hands-on IC
