AI大模型评估专家(生产力方向) - AI数据与安全
Core
Define evaluation standards and build automated assessment systems for LLM productivity and Agent capabilities to drive model optimization and product iteration.
Role type
Senior IC AI evaluation specialist (LLM productivity/Agent)
Builds
Automated evaluation frameworks (Agent-as-Judge), evaluation question banks, and high-quality assessment reports.
Domain
Artificial Intelligence / Large Language Models / Agent Systems
Deliverable
production ML models
Required skills
Prompt engineering, Python programming, Agent workflow construction, LLM evaluation methodologies, data quality analysis, multi-step task decomposition, project management.
Preferred skills
Experience in AI productivity products (coding, office, data analysis, browser agents), background in AI/CS/Data Science, experience with evaluation/annotation.
Responsibilities
Collaborate with R&D to iterate evaluation processes and standards; build and maintain automated evaluation systems; construct and update evaluation question banks; analyze end-to-end Agent task execution to attribute performance to model capabilities vs. engineering harness.
Seniority
Senior, hands-on IC
