CareerPlanSign in

Member of Technical Staff - Synthetic Data & Data Scaling

San Francisco 💼 Full-time🗓 2026-10-01 → 2026-10-02

Core

Design and run synthetic data pipelines at scale to generate, filter, weight, and scale training data for model pretraining and midtraining.

Role type

Senior IC machine-learning engineer (synthetic data & scaling)

Builds

Synthetic data generation and curation pipelines

Domain

Artificial Intelligence / Large Language Models

Deliverable

production ML models

Required skills

Synthetic data generation, data filtering, data weighting, scaling strategies, experimental design, ablation studies, model evaluation metrics

Preferred skills

Pretraining strategies, midtraining optimization, hypothesis testing, pipeline orchestration

Technologies

N/A

Responsibilities

Design and run synthetic data pipelines at scale, build methods to measure data intervention impact, run scaling and ablation experiments to determine training inputs

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 928,000+ jobs from 20+ sources.