Sr. Applied Scientist, AI Evaluation & Quality Systems
Core
Designing scalable systems and methodologies for human-in-the-loop AI evaluation, ground truth generation, and quality assurance to validate signals for training and evaluating AI/LLM systems.
Role type
Senior Applied Scientist (AI Evaluation & Quality Systems)
Builds
Production-grade evaluation pipelines, real-time monitoring systems, calibration frameworks, and root-cause analysis tooling for AI/ML systems.
Domain
Generative AI, Large Language Models (LLMs), Machine Learning Operations (MLOps), Data Quality
Deliverable
production ML models
Required skills
Ground truth generation pipeline design, real-time drift/distribution shift detection, LLM-as-a-judge evaluation methodology, calibration frameworks, root-cause analysis of judgment disagreements, Python, ML frameworks, production deployment of LLM pipelines
Preferred skills
Configurable and extensible system design, influencing technical direction across cross-functional teams, leveraging AI to improve work efficiency
Technologies
Python, LLM frameworks, ML frameworks
Responsibilities
Design and implement scalable ground truth generation pipelines across varied task types and annotation modalities; Build and maintain real-time monitoring systems for drift and quality degradation; Design calibration frameworks to re-anchor LLM evaluators against human-verified gold sets; Build root-cause analysis tooling to surface disagreement patterns between automated and human judgments; Partner with downstream users (ML teams, annotators) to ground design decisions in real feedback; Prototype, validate, and ship quality control solutions.
Seniority
Senior, hands-on IC