CareerPlanSign in

大模型推理框架研发工程师

Shanghai, China💼 Full-time🗓 2026-09-28

Core

Develop and optimize large model inference engines and PD-separated inference scheduling systems to enhance efficiency in large-scale distributed inference.

Role type

Senior IC large model inference framework engineer

Builds

High-performance distributed inference systems for large language and multimodal models

Domain

AI infrastructure / Large model inference optimization

Deliverable

production ML models

Required skills

C/C++, Python, large model inference frameworks (vLLM, SGLang, TensorRT-LLM), parallel strategies (data/pipe-line parallelism), GPU/AI chip architecture, deep learning operator implementation, system performance tuning

Preferred skills

NVLINK/GPU RDMA communication, open-source model architecture analysis, heterogeneous AI chip optimization

Sourced via tencent · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.