AI Infra研发工程师 - 存储
Core
Develop and optimize distributed parallel strategies for deep learning frameworks to support efficient large-scale model training.
Role type
Senior IC distributed systems engineer (AI training infrastructure)
Builds
Distributed training frameworks and optimization layers for large language models
Domain
AI infrastructure / Distributed systems
Deliverable
production ML models
Required skills
Deep learning framework internals, distributed parallel strategies (data/tensor/pipeline parallelism), collective communication optimization (AllReduce, AllGather), C++, Python, Linux, algorithm and data structures
Preferred skills
Storage system development, CUDA programming, GPU parallel optimization
Responsibilities
Research and optimize distributed parallel strategies for deep learning frameworks; Optimize collective communication techniques to resolve bottlenecks in large-scale clusters; Collaborate on data scheduling and caching strategies for distributed training scenarios; Drive architecture iteration and troubleshoot issues in distributed training deployments