AIML - Sr Machine Learning Engineering Manager, Evaluation
Core
Lead technical strategy and execution for agent evaluation and automatic optimization of foundation models and agentic experiences.
Role type
Senior Machine Learning Engineering Manager (hands-on IC with people leadership)
Builds
Scalable evaluation systems, benchmarks, LLM-based evaluators, simulation environments, and automated optimization pipelines for agents and foundation models.
Domain
Artificial Intelligence / Large Language Models / Agentic Systems
Deliverable
production ML models
Required skills
Python, production ML systems, large language models, agentic systems, automated evaluation methods, model refinement, synthetic data generation, research adaptation
Preferred skills
automatic prompt optimization, agent-search methods, multi-objective optimization, privacy-preserving ML, cross-organizational strategy influence
Technologies
Python, LLM frameworks, simulation environments, reward modeling, reinforcement learning
Responsibilities
Architect scalable evaluation systems including benchmarks and LLM-based judges; Establish end-to-end evaluation flywheels connecting failures to model improvements; Lead and mentor a team of ML engineers while participating in technical design and code reviews; Define technical strategy for automatic prompt, context, and agent-harness optimization; Develop methods to convert evaluation findings into actionable model-improvement signals; Partner on synthetic data generation pipelines for evaluation and post-training.
Seniority
Senior, hands-on IC with people management