CareerPlanSign in

Benchmarking Project Lead, Siri Evaluation

Cambridge, United Kingdom💼 Full-time🗓 2026-09-17 → 2026-09-28

Core

Design and execute human-in-the-loop evaluation processes to benchmark the accuracy of next-generation Siri AI across features, locales, and platforms.

Role type

Senior Manager, Evaluation & Data Science

Builds

Reliable evaluation systems and tooling ecosystems for managing rich datasets

Domain

Consumer AI / Human-in-the-loop evaluation

Deliverable

production ML models

Required skills

Agentic coding, metrics analysis, crowd science, data collection, annotation analysis, statistics, budget management, cross-functional collaboration

Preferred skills

Computational linguistics, language quality assessment, Python, data pipeline engineering

Responsibilities

Design efficient data collection processes using humans in the loop; Lead annotation efforts for various languages and devices; Plan and manage annotator resourcing budgets; Track and improve quality of human judgements via training and review mechanisms; Collaborate with engineering teams to build tooling ecosystems.

Seniority

Senior, hands-on IC with management responsibilities

Sourced via apple · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.