CareerPlanSign in

Staff Software Engineer - Observability

Sunnyvale, CA💼 Full-time🗓 2026-09-29 → 2026-10-01

Core

Design and implement metrics, logging, tracing, and alerting infrastructure to enable fast debugging and high reliability for large-scale, performance-critical distributed AI inference systems.

Role type

Staff Software Engineer (Observability & Distributed Systems)

Builds

Internal observability platforms, telemetry pipelines, libraries, and tooling for AI inference services

Domain

AI/ML Infrastructure, Distributed Systems, Observability

Deliverable

production ML models

Required skills

Backend/systems software engineering, Go/C++/Rust/Java/Python, Distributed systems, Networking fundamentals, Concurrency and performance tradeoffs, Metrics/logs/distributed tracing, Production monitoring and alerting, OpenTelemetry, Prometheus, Grafana, Datadog/Elastic/Jaeger/Tempo, High-signal alerts design, Scalable telemetry pipelines, SLIs/SLOs design

Preferred skills

High-performance computing, AI/ML systems, Inference platforms, Hardware-aware observability, SRE or platform engineering background, Large-scale production incident debugging, Internal developer platforms

Technologies

OpenTelemetry, Prometheus, Grafana, Datadog, Elastic, Jaeger, Tempo

Responsibilities

Design and implement observability instrumentation across services and platforms; Build and maintain telemetry pipelines for metrics, logs, and traces at scale; Develop internal observability platforms, libraries, and tooling; Define and operationalize SLIs, SLOs, and alerting strategies; Partner with engineers to make systems debuggable by design; Reduce MTTR by enabling fast root-cause analysis during incidents

Seniority

Staff, hands-on IC with strategic impact

Sourced via ashby · Listed on CareerPlan, which tracks 901,000+ jobs from 20+ sources.