Staff Research Engineer
Core
Define the science of model development feedback loops, establishing evaluation standards and post-training methodologies for Reddit-native LLMs.
Role type
Staff Research Engineer (Post-Training & Evaluation Science)
Builds
Foundational Large Language Models (LLMs) for Safety, Moderation, Search, and Ads.
Domain
Internet / Large Language Models / AI Safety
Deliverable
production ML models
Required skills
LLM post-training and evaluation, evaluation reliability (variance, calibration, statistical significance), custom evaluation harness engineering, generation and representation evaluation, Continuous Pre-training (CPT) and Instruction Tuning (SFT), Python, data-pipeline engineering, PyTorch, distributed training (FSDP, DeepSpeed ZeRO-3)
Preferred skills
MLflow, modern fine-tuning frameworks (Axolotl, TorchTune), synthetic data generation, preference optimization (DPO, RLHF, RLAIF, GRPO), multimodal model evaluation
Technologies
Hugging Face Transformers, vLLM, lm-eval-harness, PyTorch, FSDP2, DeepSpeed ZeRO-3, MLflow, Axolotl, TorchTune
Responsibilities
Define the 'Reddit Benchmark' evaluation standard for Safety, Reasoning, and Reddit-specific knowledge; Establish evaluation reliability and statistical rigor; Drive evaluation as a release gate in CI/CD; Design model-as-a-judge methodology; Set post-training recipes and strategy; Evaluate base and CPT checkpoints; Drive synthetic data generation strategy; Partner with Safety Engineering to translate policy into metrics; Diagnose post-training instability; Lead research direction and mentor engineers
Seniority
Staff, hands-on IC with strategic leadership