Software Engineering Director, Agentic Evaluations
Core
Lead engineering strategy and execution for credible, scalable evaluations of customer-facing AI agents, owning the core evaluation platform from APIs to production delivery.
Role type
Director-level engineering leader (AI agent evaluations)
Builds
Reusable evaluation infrastructure, APIs, workflow orchestration, and production systems for measuring agent performance, accuracy, and policy adherence.
Domain
AI/ML, Agent Systems, Software Engineering
Deliverable
production ML models | infrastructure
Required skills
Backend engineering (Python, Java/Kotlin, TypeScript/JavaScript, Go), Engineering management, LLM-as-a-judge applications, Coding-agent tools, Agent trajectory analysis, Workflow orchestration
Preferred skills
MCP servers, Benchmark frameworks (STATE-Bench, tau2-bench), Browser automation (Playwright, browser-use), Durable execution platforms (Temporal, DBOS), Agent sandboxing (AWS E2B, Daytona)
Responsibilities
Lead development of evaluations for live software agents, Own full technology stack for evaluation platform, Design reusable evaluation primitives, Turn integration processes into automated workflows, Monitor emerging evaluation frameworks, Partner with data science on proprietary benchmarks, Lead and mentor engineering team
Seniority
Director, hands-on IC with team leadership
