Data Center Global Repairs Program Support
Core
Define and manage the end-to-end hardware repair program for a global fleet of data centers, ensuring compute availability and repair SLAs.
Role type
Senior IC data center hardware operations lead
Builds
Global hardware repair program, repair SLAs, and fleet availability
Domain
Data center infrastructure / Hardware operations
Deliverable
production ML models
Required skills
Data center hardware operations management, break-fix program execution, vendor/OEM/ODM management, RMA and reverse logistics, spares planning and inventory management, failure analysis and root cause analysis, SLA definition and monitoring, process standardization, technical verification of server/network/rack-level hardware
Preferred skills
GPU/accelerator and high-density liquid-cooled infrastructure repair, hyperscale OEM/ODM RMA program management, depot repair operations, partner-operated/colocation site management, optics and high-speed interconnect failure analysis
Technologies
Server, GPU/accelerator, network, rack-level hardware, optics, high-speed interconnects
Responsibilities
Define global repair strategy including SLAs, prioritization rules, and escalation paths; Own repair turnaround time and backlog across the fleet; Author and improve procedures for triage, break-fix, and return-to-service validation; Manage RMA and reverse logistics programs with OEMs, ODMs, and vendors; Set spares pool sizing and stocking levels; Analyze failure patterns to drive corrective actions; Lead operating cadence with vendor and site leads; Communicate repair constraints and fleet availability impact to leadership
Seniority
Senior, hands-on IC