大模型推理框架研发工程师
Core
Develop and optimize large model inference engines and PD-separated inference scheduling systems to enhance efficiency in large-scale distributed inference.
Role type
Senior IC large model inference framework engineer
Builds
High-performance distributed inference systems for large language and multimodal models
Domain
AI infrastructure / Large model inference optimization
Deliverable
production ML models
Required skills
C/C++, Python, large model inference frameworks (vLLM, SGLang, TensorRT-LLM), parallel strategies (data/pipe-line parallelism), GPU/AI chip architecture, deep learning operator implementation, system performance tuning
Preferred skills
NVLINK/GPU RDMA communication, open-source model architecture analysis, heterogeneous AI chip optimization
Sourced via tencent · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.