Machine Learning Engineer, Human Centered AI - Evaluations & Insights
Core
Architect evaluation frameworks and MLOps pipelines to assess, interpret, and optimize Foundation Models and generative AI systems for human alignment.
Role type
Senior IC machine learning engineer (evaluation & insights)
Builds
Scalable evaluation suites, automated annotation pipelines, and programmatic guardrails for LLMs and multimodal models
Domain
Generative AI, Large Language Models, Human-Centered AI
Deliverable
production ML models
Required skills
Python, PyTorch, JAX, Hugging Face, LLM evaluation frameworks, RLHF/DPO, RAG architectures, embedding-based clustering, distributed inference (Ray, vLLM), MLOps automation
Preferred skills
Human factors, HCI, cognitive science methodologies
Technologies
Ray, vLLM, MLflow, Weights & Biases, G-Eval, DeepEval, SelfCheckGPT
Responsibilities
Architect comprehensive evaluation suites for LLMs and multimodal models; Develop deterministic and LLM-assisted scoring frameworks; Translate qualitative failure modes into quantifiable training signals; Build automated evaluation pipelines utilizing LLMs; Define quantitative frameworks capturing human factors like trust calibration; Collaborate cross-functionally to refine model behavior via prompt engineering and fine-tuning.
Seniority
Senior, hands-on IC
