多模态大模型数据工程师
Core
Design and develop large-scale pre-training data processing pipelines and platforms for foundation models (LLM, VLM), covering data sourcing, collection, parsing (OCR, images, web), lifecycle management, and data synthesis frameworks.
Role type
Senior IC multimodal large model data engineer
Builds
Stable, reliable data processing pipelines and platforms for base model pre-training
Domain
Artificial Intelligence / Large Language Models / Data Engineering
Deliverable
production ML models
Required skills
Python, Go, Java, Spark, Flink, Kafka, Hive, HDFS, data platform development, LLM/VLM ecosystem knowledge
Preferred skills
Data middle platform experience, machine learning system platform development, deep understanding of large model technology and product ecosystem
Technologies
Go, Python, Java, Spark, Flink, Kafka, Hive, HDFS
Responsibilities
Design and develop data processing pipelines for model pre-training; Build data platforms for metadata, lineage, and storage governance; Develop data synthesis frameworks for LLM/VLM scaling; Abstract and develop efficient data processing frameworks to boost algorithm engineer efficiency
Seniority
Senior, hands-on IC