Staff ML Platform Engineer (MLOps)
Core
Build and operate the end-to-end MLOps platform for batch, real-time, and LLM-powered workforce development models.
Role type
Staff ML Platform Engineer (MLOps)
Builds
Compute environments, model deployment pipelines, LLM routing infrastructure, and observability systems for production ML.
Domain
AI/ML infrastructure and workforce development technology.
Deliverable
production ML models
Required skills
MLOps practice ownership, LLM serving and routing, batch and real-time model operations, compute provisioning, experiment tracking, A/B testing design, production observability, data feature infrastructure.
Preferred skills
Experience with Braintrust or comparable observability tooling, background in Data Platform Engineering or Data Science.
Technologies
Braintrust, containers, CI/CD pipelines, model registries.
Responsibilities
Evaluate current ML workflows and create a prioritized platform roadmap; own compute provisioning and environment reproducibility; build LLM routing and cost-optimization layers; design and execute safe model rollouts (shadow/canary); implement drift detection and lineage tracing; debug production model failures.
Seniority
Staff, hands-on IC

