CareerPlanSign in

云监控/APM研发工程师-火山引擎

北京💼 Full-time🗓 2026-09-28

Core

Design and develop a large-scale monitoring platform for internal business and To B customers, focusing on high-throughput, low-latency, and low-cost data infrastructure.

Role type

Senior IC backend engineer (monitoring & APM)

Builds

High-scale monitoring platform, SLO management, and intelligent monitoring models (AIOps)

Domain

Cloud infrastructure, distributed systems, observability

Deliverable

production ML models

Required skills

Go/Java/Python, Linux, TCP/IP, HTTP, gRPC/Thrift/bRPC, Redis, MySQL, Kafka, RocketMQ, high-concurrency distributed system design

Preferred skills

Prometheus/OpenFalcon/Zabbix source code, VictoriaMetrics/InfluxDB/OpenTSDB, APM/distributed tracing (OpenTelemetry/SkyWalking/Jaeger), high-availability stability (disaster recovery, rate limiting, self-healing), AIOps

Technologies

Prometheus, OpenFalcon, Zabbix, VictoriaMetrics, InfluxDB, OpenTSDB, Prometheus TSDB, OpenTelemetry, SkyWalking, Jaeger, gRPC, Thrift, bRPC, Redis, MySQL, Kafka, RocketMQ

Responsibilities

Architect and develop the monitoring platform; optimize time-series data ingestion, storage, and query chains; build platform features for SLO and custom alerting; evolve intelligent monitoring models using AIOps.

Seniority

Mid-level, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.