Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)
Core
Build automated detection, diagnosis, and remediation systems for network faults in Apple's hyper-scale cloud environment to ensure fault tolerance with minimal human intervention.
Role type
Senior IC software engineer (distributed systems & fault tolerance)
Builds
Self-healing network infrastructure and automated remediation layers
Domain
Cloud networking & distributed systems
Deliverable
production ML models | product features
Required skills
distributed systems architecture, concurrency models, graph algorithms, systems design, fault-tolerant mechanisms, chaos engineering, simulation-based testing, technical ownership
Preferred skills
cloud networking architectures, SDN control planes, L3/L4 routing protocols, overlay networks, Linux networking constructs (eBPF, XDP, OVS), distributed consensus algorithms, closed-loop control systems, high-throughput telemetry pipelines
Technologies
Go, C++, Rust, Python
Responsibilities
Design and build automated detection, diagnosis, and remediation systems for network faults; Develop decision logic and safety mechanisms for automated corrective actions; Analyze failure patterns and design software solutions to mitigate systemic weaknesses; Build simulation, chaos testing, and fault-injection tooling; Apply formal reasoning about distributed systems and failure modes; Partner with production and reliability teams to understand operational constraints
Seniority
Senior, hands-on IC
