AI Interaction Evaluator (Codex
Core
Evaluate the quality of interactions between developers and modern AI coding agents (OpenAI Codex, Claude Code) to assess engineering judgment, reasoning, and output usefulness.
Role type
Senior AI Interaction Evaluator (Contract)
Builds
Evaluation frameworks and feedback loops for AI coding agents
Domain
AI-assisted software development / LLM evaluation
Deliverable
production ML models | product features
Required skills
Staff/Principal-level engineering experience, TypeScript/JavaScript, Python, hands-on experience with OpenAI Codex, Claude Code, Cursor, ability to evaluate code without full execution, opinionated feedback delivery
Preferred skills
Experience with Cursor or similar AI-first IDEs, prompt design or evaluation workflows, mentoring senior engineers, defining engineering standards
Technologies
OpenAI Codex, Claude Code, Cursor
Responsibilities
Evaluate AI-generated coding interactions end-to-end, judge output usefulness and alignment with strong engineering judgment, assess quality of explanations and reasoning, distinguish between response quality levels, provide clear feedback on what worked or felt misleading, define criteria for great interactions
Seniority
Staff / Principal-level