Observability & AgentOps Engineer
Core
Design and run the systems that keep AI agents reliable in production, focusing on monitoring, deployment, tracing, and recovery.
Role type
Senior IC Observability & AgentOps Engineer
Builds
Production-ready agent infrastructure including tracing, logging, deployment pipelines, and recovery systems.
Domain
Financial Services & Insurance / AI Agent Operations
Deliverable
production ML models
Required skills
distributed tracing, metrics, structured logging, SLIs/SLOs definition, alerting design, CI/CD, Docker, AWS Step Functions, Python 3.11+
Preferred skills
LLM/agent workload design, infrastructure-as-code, high-volume log/trace storage design
Technologies
Langfuse, OpenTelemetry, ELK, AWS Step Functions, Docker, Python
Responsibilities
Design end-to-end observability architecture (tracing, metrics, logging); Define reliability targets (SLIs/SLOs) and alerting; Design deployment topology and orchestration; Design failure recovery mechanisms; Design log and trace data model and storage; Set up day-to-day monitoring and support handover.