CareerPlanSign in

Member of Technical Staff, Evals

San Francisco, CA💼 Full-time🗓 2026-09-29 → 2026-09-30

Core

Design, build, and publish coding benchmarks and evaluation environments to measure and improve frontier coding agents.

Role type

Senior IC machine-learning engineer (coding evaluations)

Builds

Coding benchmarks, evaluation environments, verifiers, graders, and feedback systems for agentic software development

Domain

AI / Software Engineering / Evaluation

Deliverable

production ML models

Required skills

Python, software engineering, experimental design, system design, hypothesis testing, failure analysis, open-source contribution

Preferred skills

Coding benchmark creation, automated graders, reinforcement learning, post-training methods, research publication

Technologies

Python, modern ML tooling, evaluation infrastructure, data workflows

Responsibilities

Design and publish coding benchmarks that challenge state-of-the-art agents; Create realistic software-engineering tasks and test harnesses; Develop verifiers and reward signals for agentic software development; Analyze coding-agent behavior to diagnose failure modes; Partner with researchers to develop high-signal evaluation methods; Productize repeatable patterns into reusable software and platforms

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 880,000+ jobs from 20+ sources.