Engineering Manager, Cloud Monitoring Services Platform
Core
Lead the Platform team for Crusoe Cloud's observability stack, managing time series and log storage systems and the query layer that serves dashboards and APIs for AI workloads.
Role type
First-line Engineering Manager (Platform)
Builds
Time series and log storage systems, query layer, and telemetry retention infrastructure for AI training and inference fleets.
Domain
Cloud Infrastructure / AI Workloads / Observability
Required skills
Backend or distributed systems engineering, team management (4-6 engineers), technical direction for storage/query, operational incident management, cross-functional coordination, hiring and onboarding
Preferred skills
Time series databases (Prometheus, Thanos, VictoriaMetrics, InfluxDB, ClickHouse), query engine optimization, log storage (Loki, Elasticsearch), multi-tenant system design, GPU/accelerated computing environments
Technologies
Go, Rust, Java, C++, Kubernetes, Prometheus, Thanos, Mimir, VictoriaMetrics, InfluxDB, ClickHouse, Loki, Elasticsearch, OpenSearch, OpenTelemetry, Grafana
Responsibilities
Manage and grow a team of 4-6 engineers, own storage and query system performance/cost, plan and sequence delivery roadmap, set technical direction alongside Staff engineers, maintain high operational bar for critical path queries, manage oncall rotation, coordinate dependencies with collection/ingestion teams, shape team through recruiting and onboarding
Seniority
Manager, hands-on IC leadership