CareerPlanSign in

AI大模型评估专家(生产力方向) - AI数据与安全

北京💼 Full-time🗓 2026-09-28

Core

Define evaluation standards and build automated assessment systems for LLM productivity and Agent capabilities to drive model optimization and product iteration.

Role type

Senior IC AI evaluation specialist (LLM productivity/Agent)

Builds

Automated evaluation frameworks (Agent-as-Judge), evaluation question banks, and high-quality assessment reports.

Domain

Artificial Intelligence / Large Language Models / Agent Systems

Deliverable

production ML models

Required skills

Prompt engineering, Python programming, Agent workflow construction, LLM evaluation methodologies, data quality analysis, multi-step task decomposition, project management.

Preferred skills

Experience in AI productivity products (coding, office, data analysis, browser agents), background in AI/CS/Data Science, experience with evaluation/annotation.

Responsibilities

Collaborate with R&D to iterate evaluation processes and standards; build and maintain automated evaluation systems; construct and update evaluation question banks; analyze end-to-end Agent task execution to attribute performance to model capabilities vs. engineering harness.

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.