CareerPlanSign in

Staff Research Engineer

USA🌐 Remote💼 Full-time💰 $230,000–$230,000🗓 2026-07-14 → 2026-09-29

Core

Define the science of model development feedback loops, establishing evaluation standards and post-training methodologies for Reddit-native LLMs.

Role type

Staff Research Engineer (Post-Training & Evaluation Science)

Builds

Foundational Large Language Models (LLMs) for Safety, Moderation, Search, and Ads.

Domain

Internet / Large Language Models / AI Safety

Deliverable

production ML models

Required skills

LLM post-training and evaluation, evaluation reliability (variance, calibration, statistical significance), custom evaluation harness engineering, generation and representation evaluation, Continuous Pre-training (CPT) and Instruction Tuning (SFT), Python, data-pipeline engineering, PyTorch, distributed training (FSDP, DeepSpeed ZeRO-3)

Preferred skills

MLflow, modern fine-tuning frameworks (Axolotl, TorchTune), synthetic data generation, preference optimization (DPO, RLHF, RLAIF, GRPO), multimodal model evaluation

Technologies

Hugging Face Transformers, vLLM, lm-eval-harness, PyTorch, FSDP2, DeepSpeed ZeRO-3, MLflow, Axolotl, TorchTune

Responsibilities

Define the 'Reddit Benchmark' evaluation standard for Safety, Reasoning, and Reddit-specific knowledge; Establish evaluation reliability and statistical rigor; Drive evaluation as a release gate in CI/CD; Design model-as-a-judge methodology; Set post-training recipes and strategy; Evaluate base and CPT checkpoints; Drive synthetic data generation strategy; Partner with Safety Engineering to translate policy into metrics; Diagnose post-training instability; Lead research direction and mentor engineers

Seniority

Staff, hands-on IC with strategic leadership

Sourced via codingjobboard · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.