Code大模型工程师(数据方向)-搜索
Core
Design and develop data processing pipelines and frameworks for training Code Large Language Models, focusing on data production, management, and scaling.
Role type
Senior IC machine-learning data engineer (code LLM)
Builds
Pretrain data pipelines and data processing engines for Code LLMs
Domain
Artificial Intelligence / Large Language Models / Data Engineering
Deliverable
production ML models
Required skills
Python, Go, Java, Spark, Flink, Kafka, Hive, HDFS, data platform development, LLM pretraining, data synthesis
Preferred skills
Agent-based systems, Self-Evolution techniques, data middle platform experience
Responsibilities
Design and develop multi-type data processing pipelines for LLM training; Build data synthesis solutions to support data scaling; Abstract and develop efficient, reliable data processing frameworks; Explore Agent and Self-Evolution methods for next-gen Pretrain data pipelines.