CareerPlanSign in

大模型预训练工程师-AI Data

上海💼 Full-time🗓 2026-09-28

Core

Build end-to-end pipelines for sourcing, collecting, parsing, and processing massive-scale high-quality pre-training data for large language models (LLMs), including research on data synthesis and automated evaluation.

Role type

Senior IC LLM pre-training data engineer (synthesis & pipeline)

Builds

Automated data production pipelines, data synthesis frameworks, and evaluation systems for LLM pre-training

Domain

Artificial Intelligence / Large Language Models / Data Engineering

Deliverable

production ML models

Required skills

Python, distributed computing frameworks (Spark, Flink, Ray), LLM architecture and training mechanisms, data synthesis techniques (Self-Instruct, Agent simulation), automated evaluation (Reward Models, LLM-as-a-Judge), data quality metrics design, machine learning algorithms for data filtering

Preferred skills

Experience with RLHF/DPO data construction, vLLM, open-source contributions in NLP/LLM, publications in top-tier conferences (ACL, EMNLP, NeurIPS)

Responsibilities

Design and optimize data engineering infrastructure for massive data cleaning, deduplication, and formatting; develop strategies to supplement data gaps using LLM-based synthesis; establish automated evaluation systems to iterate on data generation strategies; collaborate with algorithm and infrastructure teams to optimize real vs. synthetic data ratios

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.