CareerPlanSign in

Data Center Hardware Quality & Reliability Engineer

San Francisco💼 Full-time🗓 2026-09-28 → 2026-10-02

Core

Own the end-to-end data-center hardware quality and reliability loop for OpenAI's 3P infrastructure and 1P current and next-gen platforms, turning field failures into quantified risk and upstream changes.

Role type

Senior IC hardware reliability engineer (data center)

Builds

Data-center hardware systems (servers, racks, infrastructure)

Domain

Data center infrastructure / Hardware reliability

Required skills

Hardware/system architecture, Reliability statistics (Weibull/Poisson/MTBF), FMEA/FTA, 8D/CAPA, Root cause analysis, SQL, Python/R

Preferred skills

GPU/AI server platforms, Liquid cooling, High-power delivery, Linux/BMC/IPMI/Redfish logs, ODM/CM supplier experience, Leadership of cross-generation reliability programs

Technologies

SQL, Python, R, Linux, BMC, IPMI, Redfish

Responsibilities

Build and govern field-quality data models; Define reliability metrics (AFR, MTBF, etc.); Lead systemic field-failure triage and failure analysis; Develop spare-demand projections; Partner with MQE/NPI to convert field mechanisms into manufacturing-test coverage; Run the cross-functional reliability council.

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via ashby · Listed on CareerPlan, which tracks 924,000+ jobs from 20+ sources.