Senior Software Development Engineer in Test — LLM Evaluation & Automation, T3E
Core
Design and implement automated LLM evaluation harnesses to catch model regressions in generative AI features before they reach human evaluation or live rollout.
Role type
Senior SDET (LLM Evaluation & Automation)
Builds
Automated evaluation pipelines, reusable test libraries, and quality gate reporting for generative AI features.
Domain
Generative AI / Large Language Models / Test Automation
Deliverable
production ML models
Required skills
Test automation, Python, LLM-as-a-judge evaluation, CI/CD integration, software engineering fundamentals, debugging and triage, rubric design, data pipeline fluency
Preferred skills
On-device tooling, generative model behavior (image/NLP), dataset curation, SQL, Xcode, fairness considerations
Technologies
Python, JSON, YAML, REST APIs, SQL, Jupyter, CI/CD systems
Responsibilities
Design and maintain LLM-as-a-judge evaluation harnesses integrated into CI/automation pipelines; Author and curate eval sets and rubrics; Run eval jobs at scale and triage results to distinguish regressions from noise; Plan and build test coverage from component tests to end-to-end user flows; Package eval tooling as reusable libraries; Define quality gates and reporting for model regressions; Collaborate with data scientists on regression detection across large output sets.
Seniority
Senior, hands-on IC