Research Engineer - Environments, Data and Post-Training
Core
Design and operate post-training and RLVR pipelines, synthetic data generation, and large-scale evaluation workflows to train frontier language models for tool use, agentic behavior, and real-world reasoning.
Role type
Senior IC research engineer (post-training & evaluation)
Builds
Scalable data augmentation pipelines, reward-shaping experiments, rubrics, evaluators, and LLM benchmark systems
Domain
AI/ML, Large Language Models, Post-training, RLVR, Synthetic Data
Deliverable
production ML models
Required skills
Applied research in post-training/model evaluation, Machine learning model development, Backend systems engineering, Data structures and algorithms, API integration, SQL/NoSQL databases, Cloud platforms, Experimental design and analysis, Data quality assessment
Preferred skills
Real-world post-training team experience, Publications at top-tier conferences (NeurIPS, ICML, ACL), Synthetic data generation, RL-style workflows
Technologies
GRPO, DAPO, LLM evaluation frameworks, Cloud platforms
Responsibilities
Work on post-training and RLVR pipelines to optimize datasets, rewards, and training strategies; Design and run reward-shaping experiments and algorithmic improvements; Quantify data usability, quality, and performance uplift on benchmarks; Build and maintain scalable data generation and augmentation pipelines; Create and refine rubrics, evaluators, and scoring frameworks; Build and operate LLM evaluation systems, benchmarks, and metrics at scale
Seniority
Senior, hands-on IC