Agent效果评测专家 - 开发者服务
Core
Define evaluation standards and acceptance criteria for AI Agents in software engineering scenarios, build high-fidelity test sets, and design quantitative metrics to systematically evaluate Agent performance.
Role type
Senior AI Agent Evaluation Engineer
Builds
High-fidelity test sets, quantitative evaluation metrics, and automated evaluation pipelines for AI Agents
Domain
Software Engineering / AI Agents / Large Language Models
Deliverable
production ML models
Required skills
Go, Python, Java, AI Agent architecture (Multi-Agent, Context Engineering, ReAct), LLM evaluation frameworks, root cause analysis, automated testing
Preferred skills
Agent development experience, complex scenario evaluation, AI paper publication, LLM training experience
Technologies
Go, Python, Java, MCP, A2A, Function Call
Responsibilities
Define evaluation standards and acceptance criteria for AI Agents in software engineering scenarios; Build high-fidelity test sets and design quantitative metrics; Conduct root cause analysis and collaborate with PMs and algorithms to optimize Agent performance; Build automated evaluation and insight analysis capabilities; Track industry trends and validate new technologies in business scenarios.