Fellowship : (Agent Intelligence & Evaluations)
Core
Design audio-native evaluation frameworks and observability systems for self-healing voice agents to diagnose failures in ASR, LLM reasoning, and tool execution.
Role type
Research Engineer (Voice AI & Agent Evaluation)
Builds
Evaluation datasets, LLM-as-judge rubrics, and end-to-end observability schemas for voice agents
Domain
Voice AI, Large Language Models, Agent Systems
Deliverable
production ML models
Required skills
Python, speech models (Whisper, Conformer), LLM tool-use frameworks, observability stacks (OpenTelemetry, Langfuse, Arize, Hamming), adversarial dataset generation, LLM-as-judge rubrics
Preferred skills
evaluation methodology, dataset synthesis, interpretability, real-time systems, telephony, streaming pipelines
Technologies
Whisper, Conformer, OpenTelemetry, Langfuse, Arize, Hamming, Python
Responsibilities
Design audio-native metrics for barge-in, prosody, and latency-induced errors; generate adversarial conversational datasets; build LLM-as-judge rubrics for task completion and empathy; shape schemas for correlating audio packets, STT hypotheses, and tool calls; mine production traces for failure patterns to drive fine-tuning or prompt updates
Seniority
PhD student (preferred) or exceptional MS student/Research Engineer