Member of Technical Staff, Evals
Core
Design, build, and publish coding benchmarks and evaluation environments to measure and improve frontier coding agents.
Role type
Senior IC machine-learning engineer (coding evaluations)
Builds
Coding benchmarks, evaluation environments, verifiers, graders, and feedback systems for agentic software development
Domain
AI / Software Engineering / Evaluation
Deliverable
production ML models
Required skills
Python, software engineering, experimental design, system design, hypothesis testing, failure analysis, open-source contribution
Preferred skills
Coding benchmark creation, automated graders, reinforcement learning, post-training methods, research publication
Technologies
Python, modern ML tooling, evaluation infrastructure, data workflows
Responsibilities
Design and publish coding benchmarks that challenge state-of-the-art agents; Create realistic software-engineering tasks and test harnesses; Develop verifiers and reward signals for agentic software development; Analyze coding-agent behavior to diagnose failure modes; Partner with researchers to develop high-signal evaluation methods; Productize repeatable patterns into reusable software and platforms
Seniority
Senior, hands-on IC
