Machine Learning Engineer - AI Evaluation & LLM Systems
Core
Build scalable infrastructure and intelligent evaluators to measure and improve the quality of large language models and multimodal AI systems.
Role type
Machine Learning Engineer (AI Evaluation & LLM Systems)
Builds
Evaluation systems, benchmarking frameworks, and scalable infrastructure for AI quality
Domain
Artificial Intelligence, Large Language Models, Multimodal AI
Deliverable
production ML models
Required skills
Python, C++, PyTorch, TensorFlow, JAX, supervised learning, statistical analysis, data processing, model training, experimentation
Preferred skills
LLMs, multimodal AI, generative AI, Git, CI/CD, distributed computing, cloud platforms, large-scale data processing
Technologies
PyTorch, TensorFlow, JAX, Git, CI/CD
Responsibilities
Design and deploy evaluation systems and scalable software; build infrastructure for model training and benchmarking; analyze model performance and identify quality issues; collaborate with researchers and product teams to translate research into production solutions; contribute to architecture and engineering best practices
Seniority
Junior to Mid-level, hands-on IC