Research Engineer, Evals - Member of Technical Staff
Core
Designing rigorous benchmarks, evaluation suites, and methodologies to analyze and evaluate agentic systems, including multi-turn, long-horizon, and tool-using behaviors.
Role type
Research Engineer, Evals (Member of Technical Staff)
Builds
Evaluation infrastructure, benchmarks, failure taxonomies, and observability tools for agentic systems
Domain
AI/ML, Agentic Systems, Heterogeneous Compute
Deliverable
production ML models
Required skills
LLMs in agentic settings, experimental design, statistical analysis, software engineering, failure analysis, automated grading
Preferred skills
Published research in ML/benchmarks, Bayesian inference, model selection, automated grading failure modes
Technologies
LLMs, agentic frameworks, statistical tools
Responsibilities
Design benchmarks for agentic behavior, build evaluation methodology, red-team evaluations, turn traces into structured evidence, infer system capabilities from partial evidence, build evaluation infrastructure
Seniority
Mid-Senior, hands-on IC