CareerPlanSign in

ASKUSR0147035 Site Reliability Engineer - NERSC Operations

Berkeley, California, United States💼 Full-time💰 $75–$75🗓 2026-09-29 → 2026-09-30

Core

Monitor and maintain high-performance computing (HPC) and data center systems for U.S. Department of Energy researchers, ensuring 24/7 availability and reliability.

Role type

Site Reliability Engineer (HPC Operations)

Builds

Reliable HPC infrastructure and data services for scientific research

Domain

High Performance Computing / Data Center Operations

Required skills

Linux shell, scripting (C/C++/Perl/Java/Python), network protocols, network security, incident triage, automation, physical data center inspection

Preferred skills

Kubernetes, Prometheus, VictoriaMetrics, Alertmanager, building management software, evaporative cooling systems, power utilization monitoring, ServiceNow, agentic AI tools

Technologies

Linux, SSH, ServiceNow, HPC systems, cooling infrastructure

Responsibilities

Monitor HPC, storage, network, and facility systems during overnight shifts; maintain operational data collection and document incidents; develop monitoring tools and alerting integrations; automate routine responses and prevent recurring incidents; conduct physical and logical data center walkthroughs; coordinate with teams on incidents and maintenance workflows.

Seniority

Mid-level, hands-on IC

Sourced via workable · Listed on CareerPlan, which tracks 862,000+ jobs from 20+ sources.