混元评测Infra可观测性工程师(北京/深圳)
Core
Build end-to-end observability for a large-scale model evaluation platform, ensuring stability and rapid fault resolution across task scheduling, model invocation, and result generation.
Role type
Senior IC observability engineer (distributed systems)
Builds
Production-grade evaluation platform infrastructure and diagnostic tooling
Domain
AI/ML infrastructure + distributed systems
Deliverable
infrastructure
Required skills
Linux troubleshooting, Kubernetes, Python, Go, distributed systems fault diagnosis, log analysis, alerting system design, automation development
Preferred skills
Large model log analysis, anomaly clustering, root cause analysis automation
Responsibilities
Monitor platform stability and respond to online faults; troubleshoot user issues in task execution and model calling; design and maintain observability dashboards and alerting rules; develop automated diagnostic tools and standardize troubleshooting workflows.