Agent评测专家(To B方向) - AI数据与安全
Core
Design and execute evaluation frameworks for AI Agents and Coding models, transforming abstract capability requirements into observable test cases and scalable data production pipelines.
Role type
Senior AI Model Evaluation Engineer (Agent/Coding)
Builds
Scalable evaluation datasets, automated assessment pipelines, and quality metrics for AI Agents and Coding models.
Domain
Artificial Intelligence / Large Language Models / Agent Systems
Deliverable
production ML models
Required skills
LLM/Agent architecture understanding, Python programming, test case design, data pipeline engineering, project management, root cause analysis
Preferred skills
Experience with automated evaluation tools, ability to define expert personas for data production, proficiency in analyzing model failure modes
Technologies
Python, LLMs, Agent frameworks, Evaluation platforms
Responsibilities
Define executable evaluation schemes for Agent and Coding models; Organize external experts to produce standardized evaluation data; Engineer automated assessment workflows and integrate them into internal platforms; Collaborate with research and product teams to drive model iteration based on evaluation data.
