混元大模型平台研发工程师(北京/深圳)
Core
Designing and developing the architecture for the Hunyuan machine learning research foundation and LLMOps engineering platform to support large model training, inference, evaluation, and data processing.
Role type
Senior IC machine-learning platform engineer (LLMOps)
Builds
Large-scale model training infrastructure, distributed training systems, and model serving platforms
Domain
Artificial Intelligence / Large Language Models / Cloud Infrastructure
Deliverable
production ML models
Required skills
PyTorch, TensorFlow, DeepSpeed, Python, Go, Java, distributed training, model training optimization, model serving optimization, software architecture design, complex system development
Preferred skills
Experience with trillion-parameter model training and deployment, AIGC engineering, LLMOps体系建设
Technologies
PyTorch, TensorFlow, DeepSpeed, Python, Go, Java
Responsibilities
Architecting and developing the core Hunyuan platform for large model training and inference; Iterating on the LLMOps engineering system to optimize platform usability, automation, and standardization; Solving large-scale training and data processing challenges to improve stability and performance.
Seniority
Senior, hands-on IC