AI评测专家 - 飞书
Core
Design and maintain the evaluation system for AI Agents, covering capabilities, business impact, tool usage, task execution, stability, safety, and user experience.
Role type
Senior AI Agent Evaluation Engineer
Builds
Automated evaluation pipelines, high-quality evaluation datasets, and failure analysis systems for AI Agents
Domain
Enterprise AI / Large Language Models / AI Agents
Deliverable
production ML models
Required skills
AI Agent evaluation, LLM evaluation, data analysis, test development, Python, SQL, LLM-as-a-Judge, Prompt Engineering, RAG, tool calling evaluation, Computer Use Agent evaluation
Preferred skills
Automated evaluation platform construction, complex task E2E evaluation, failure case root cause analysis, industry benchmark tracking
Technologies
Python, SQL, LLM-as-a-Judge
Responsibilities
Define evaluation standards for AI Agents across multiple dimensions; Design evaluation sets, scoring rules, and acceptance thresholds; Build and maintain automated evaluation capabilities and datasets; Establish failure case attribution systems for hallucinations and execution errors; Correlate offline metrics with online user feedback to refine evaluation methodologies.
