CareerPlanSign in

多模态大模型数据工程师-产品研发

上海💼 Full-time🗓 2026-09-28

Core

Design and develop large-scale pre-training data processing pipelines and platforms for foundation models (LLM, VLM), including data sourcing, scraping, parsing, lifecycle management, and synthetic data generation.

Role type

Senior IC multimodal large model data engineer

Builds

Stable, reliable data processing pipelines and platforms for base model pre-training

Domain

Artificial Intelligence / Large Language Models / Multimodal AI

Deliverable

production ML models

Required skills

Python, Go, Java, Spark, Flink, Kafka, Hive, HDFS, data pipeline architecture, data governance, synthetic data generation frameworks

Preferred skills

Data middle platform development, machine learning system platform development, deep understanding of LLM/VLM ecosystem, data experimentation engineering

Responsibilities

Design and develop data processing pipelines for model pre-training; Build data platforms for metadata, lineage, and storage governance; Develop data synthesis solutions and frameworks for scaling; Abstract and develop efficient data processing frameworks for algorithm engineers

Sourced via bytedance · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.