Senior Software Engineer, Chaos Engineering
Core
Build automation for zonal resilience, fault injection, and incident replay to ensure production systems can safely evacuate and recover from failures.
Role type
Senior Software Engineer (Chaos Engineering & Distributed Systems)
Builds
Zonal-resilience automation, fault-injection systems, gRPC services, Kubernetes controllers, and shared platform components.
Domain
Cloud Infrastructure / Observability / Distributed Systems
Required skills
Distributed systems fundamentals, Kubernetes workload lifecycles, gRPC, fault injection design, blast-radius controls, incident replay orchestration, gameday execution, technical design, mentoring
Preferred skills
Reliability engineering, chaos engineering, zonal failover, AI-assisted operational workflows, traffic interception, large-scale observability systems
Responsibilities
Build zonal-resilience automation for safe workload evacuations and recovery; Design and build fault-injection systems for production environments; Develop safeguards like kill switches and rollback paths; Build agents to propose failure scenarios and triage results; Lead gamedays from hypothesis to validation; Design and implement reliable distributed systems including Kubernetes controllers.
Seniority
Senior, hands-on IC
