混元大模型中台研发负责人(北京/深圳)
Core
Designing and implementing the technical architecture for a large model middle platform covering data, pre-training, post-training, and evaluation pipelines.
Role type
Principal Machine Learning Platform Engineer (Large Model Infrastructure)
Builds
Scalable distributed training infrastructure, data pipelines, and evaluation engines for large language models.
Domain
Artificial Intelligence / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
Kubernetes, Volcano, Ray, Spark, PyTorch, Megatron-LM, FSDP, GPU cluster management, distributed storage systems, parallel training frameworks, model evaluation engineering.
Preferred skills
Experience with Ceph, HDFS, DP/TP/PP/EP parallelism, RLHF pipelines, multi-modal model evaluation.
Responsibilities
Define technical architecture for data, pre-training, post-training, and evaluation chains; build data foundations including pipelines and metadata management; optimize distributed storage and compute systems; ensure performance and stability of large-scale distributed training; platformize training and post-training engineering workflows; manage model evaluation engineering chains; establish technical standards and release processes.
Seniority
Senior, hands-on IC with team leadership