Senior Research Scientist, Agent Evaluation
Core
Designing evaluation methodologies and benchmarks to validate AI agents before they reach customers, including LLM-as-a-Judge and trajectory-based approaches.
Role type
Senior Research Scientist (AI Agent Evaluation)
Builds
Automated evaluation pipelines and scoring logic for agent traces
Domain
Applied AI, Agentic AI, LLM Evaluation
Deliverable
production ML models
Required skills
Python, Agentic AI, Context Engineering, Prompt Engineering, LLM APIs, Evaluation Methodology Design, Synthetic Data Curation
Preferred skills
AI-assisted development tools (Claude code, Windsurf)
Technologies
LLM APIs, Python
Responsibilities
Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety; Build and harden pipelines that run traces through model-based judges at scale; Curate synthetic and real-world datasets; Measure the evaluator itself for consistency and agreement with human labels
Seniority
Senior, hands-on IC
