AI Evaluations Engineer, US Decision Intelligence
Core
Own the end-to-end evaluation pipeline for AI products and agentic workflows, defining standards and pass/fail contracts to ensure quality before release.
Role type
Senior IC AI Evaluations Engineer
Builds
Evaluation frameworks, instrumentation, and release gates for LLM and agent systems
Domain
AI/ML, Decision Intelligence, Sales Technology
Deliverable
production ML models
Required skills
Python, AI evaluation techniques (Golden datasets, LLM-as-a-Judge, rubric-based scoring), LLM ecosystems (OpenAI, Anthropic, Gemini), RAG pipelines, vector databases, SQL, CI/CD workflows, LLM observability tools
Preferred skills
Embeddings, retrieval algorithms, distributed systems architecture, asynchronous messaging, GenAI strategies, agent evaluation concepts
Technologies
Pinecone, FAISS, Milvus, PostgreSQL, Langfuse, RabbitMQ, Redis, Valkey, Hadoop, Spark, Snowflake
Responsibilities
Architect comprehensive Evals framework to trace agent responses and tool calling; Build and operate AI evaluation workflows for LLM outputs; Implement rubric-based evals for correctness and grounding; Move beyond LLM-as-a-judge to agent-as-a-judge patterns; Instrument LLM and agent workflows to capture traces and metadata; Own the platform-wide eval gate and hold releases; Define rerun policy and variance thresholds; Partner with AI engineers to translate product requirements into evaluation criteria
Seniority
Senior, hands-on IC
