Senior Research Scientist, Agentic AI
Core
Design and build automated evaluation platforms that validate AI agents using model-based judges and scoring logic to ensure reliability and safety before customer deployment.
Role type
Senior Research Scientist (Agentic AI Evaluation)
Builds
Automated evaluation pipelines, model-based judge systems, and scoring logic for AI agents
Domain
Artificial Intelligence, Agentic AI, Evaluation Methodologies
Deliverable
production ML models
Required skills
Agentic AI, LLM-as-a-Judge, Prompt Engineering, Python, Context Engineering, Evaluation Methodology Design, Synthetic Data Curation, LLM API Integration
Preferred skills
AI-assisted development tools (Claude code, Windsurf)
Technologies
Python, LLM APIs
Responsibilities
Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety; Build and harden pipelines that run traces through model-based judges at scale; Curate synthetic and real-world datasets; Measure evaluator consistency and agreement with human labels
Seniority
Senior, hands-on IC
