大模型推理架构研发工程师(J95970)
Core
Design, develop, and optimize inference frameworks and high-performance computing libraries for large language models (LLMs) and deep learning tasks.
Role type
Senior IC machine-learning infrastructure engineer (LLM inference)
Builds
PaddlePaddle inference framework, high-performance computing (HPC) libraries, and communication libraries
Domain
Artificial Intelligence / Deep Learning / Large Language Models
Deliverable
production ML models
Required skills
C++, Python, CUDA programming, computer architecture, assembly-level development, LLM core technologies (FlashAttention, PagedAttention, MoE, Chunked Prefill), quantization algorithms (AWQ, GPTQ, SmoothQuant), communication operators (Allreduce), compute-communication overlap, separated deployment (PD separation)
Preferred skills
PaddlePaddle, PyTorch, TensorFlow, vLLM, TGI, SGLang, TensorRT-LLM
Responsibilities
Optimize inference performance for Baidu's ERNIE Bot and other business LLMs; design and develop PaddlePaddle inference framework; perform deep optimization of functional modules on CPU/GPU; track and implement breakthroughs in deep learning inference technologies; optimize framework usability for developers; design and develop HPC platforms and libraries.
Seniority
Senior, hands-on IC