ASKUSR0147035 Site Reliability Engineer - NERSC Operations
Core
Monitor and maintain high-performance computing (HPC) and data center systems for U.S. Department of Energy researchers, ensuring 24/7 availability and reliability.
Role type
Site Reliability Engineer (HPC Operations)
Builds
Reliable HPC infrastructure and data services for scientific research
Domain
High Performance Computing / Data Center Operations
Required skills
Linux shell, scripting (C/C++/Perl/Java/Python), network protocols, network security, incident triage, automation, physical data center inspection
Preferred skills
Kubernetes, Prometheus, VictoriaMetrics, Alertmanager, building management software, evaporative cooling systems, power utilization monitoring, ServiceNow, agentic AI tools
Technologies
Linux, SSH, ServiceNow, HPC systems, cooling infrastructure
Responsibilities
Monitor HPC, storage, network, and facility systems during overnight shifts; maintain operational data collection and document incidents; develop monitoring tools and alerting integrations; automate routine responses and prevent recurring incidents; conduct physical and logical data center walkthroughs; coordinate with teams on incidents and maintenance workflows.
Seniority
Mid-level, hands-on IC