CareerPlanSign in

Reliability & Observability Analyst I

Sydney, New South Wales💼 Full-time🗓 2026-09-27 → 2026-09-29

Core

Analyze operational signals, improve incident quality, and support AIOps-enabled automation for 24/7 HPC Data Center Operations.

Role type

Entry-level Reliability & Observability Analyst (Level 1)

Builds

Operational reliability and detection quality for high-performance compute infrastructure

Domain

Data Center Operations / High-Performance Computing

Deliverable

dashboards & analysis

Required skills

Incident analysis, signal validation, alert quality assessment, metrics/log analysis, Linux systems, networking basics, observability platforms (Splunk, Datadog, Prometheus), AIOps concepts, automation artifact reading, trend analysis

Preferred skills

SRE concepts (MTTR/MTTD), cross-functional collaboration

Technologies

Splunk, Datadog, Prometheus, Python, Bash

Responsibilities

Analyze incidents and validate operational signals, assess alert quality and identify monitoring gaps, review and validate automated insights from AIOps tools, analyze incident trends and system behaviors, assist with minor updates to automation artifacts under guidance

Seniority

Entry-level, hands-on IC

Sourced via viewjobs · Listed on CareerPlan, which tracks 847,000+ jobs from 20+ sources.