Sr. Staff Engineer (Reliability Engineering)
Core
Architect and lead enterprise critical platforms to prevent, manage, and resolve high-severity technology incidents, ensuring fault-tolerant infrastructure and minimizing downtime.
Role type
Sr. Staff Reliability Engineer (Strategic IC + Mentorship)
Builds
Enterprise critical platforms, microservice architectures, event-driven platforms, and robust observability/disaster recovery systems
Domain
Banking / Cloud Infrastructure / Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Software engineering and solution architecture, Enterprise architecture and design patterns, Cloud computing (AWS, Azure, GCP), Data architecture, Distributed systems design, Mentoring senior engineers, Cross-domain architectural leadership, Agentic AI coding tools
Preferred skills
Master's in CS/SE, 12+ years in distributed HPC and ML systems, 12+ years coding in Go/Java/Python/Rust/C#/Scala, 10+ years with database management systems, Networking protocols and third-party distributed systems integration
Technologies
AWS, Microsoft Azure, Google Cloud, Go, Java, JavaScript, TypeScript, Python, Rust, C#, Scala, Kubernetes, Docker, Terraform, Prometheus, Grafana, Chaos Engineering tools
Responsibilities
Provide architectural vision and lead cross-functional teams during critical outages, spearhead systemic root cause analyses, define enterprise-wide standards for observability and disaster recovery, drive modernization of monolithic systems into microservices, mentor engineers and coach technical teams, collaborate on long-term platform strategy and API standards, act as a trusted advisor for key technologies
Seniority
Sr. Staff, hands-on IC with strategic leadership and mentorship

