CareerPlanSign in

生态发展部-资深游戏测试工程师-Agent评测

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Design and execute full-link quality assurance and benchmarking for AI Agents in game generation, focusing on establishing automated evaluation baselines and defining graded metrics.

Role type

Senior IC AI Agent evaluation engineer (game generation)

Builds

Automated evaluation baselines, high-quality benchmark datasets, and multi-dimensional evaluation reports

Domain

AI Agents, Game Development, Large Language Models

Deliverable

production ML models

Required skills

AI Agent evaluation, Benchmark construction, LLM-as-a-Judge, Root cause analysis, Python programming, Automated testing scripts

Preferred skills

Game generation business logic understanding, Cross-team collaboration

Technologies

LLM-as-a-Judge, Automated testing frameworks, Log analysis tools, Traceability tools

Responsibilities

Design and quantify evaluation standards for non-deterministic AI outputs; Build benchmark datasets covering multi-turn interactions and long/short contexts; Develop automated evaluation scripts and integrate LLM-based verification; Establish defect triage mechanisms to identify root causes; Collaborate with development teams to resolve issues based on evaluation reports.

Sourced via tencent · Listed on CareerPlan, which tracks 852,000+ jobs from 20+ sources.