大模型训练优化工程师 - Seed Model
Core
Design and develop large-scale machine learning system architectures to optimize training efficiency and reliability for foundational AI models.
Role type
Senior IC machine learning systems engineer (LLM training optimization)
Builds
Scalable distributed training infrastructure for foundational models supporting applications like Doubao and Jiemeng
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
Distributed model training, High-performance computing, CUDA programming, System architecture design, Resource scheduling, Data management, Algorithm-system co-optimization, Pretrain/RL training, New hardware adaptation
Preferred skills
LLM/NLP/CV/voice algorithms, Diffusion/RL algorithms, Torch.compile/Triton/TVM, RDMA communication libraries, Heterogeneous accelerated hardware, Big data architecture
Responsibilities
Design scalable ML system architectures, Research and implement cutting-edge training technologies, Collaborate with algorithm teams on joint optimization, Manage distributed training and resource scheduling
Seniority
Senior, hands-on IC