CareerPlanSign in

大模型推理架构研发工程师(J95970)

北京市,上海市,深圳市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Design, develop, and optimize inference frameworks and high-performance computing libraries for large language models (LLMs) and deep learning tasks.

Role type

Senior IC machine-learning infrastructure engineer (LLM inference)

Builds

PaddlePaddle inference framework, high-performance computing (HPC) libraries, and communication libraries

Domain

Artificial Intelligence / Deep Learning / Large Language Models

Deliverable

production ML models

Required skills

C++, Python, CUDA programming, computer architecture, assembly-level development, LLM core technologies (FlashAttention, PagedAttention, MoE, Chunked Prefill), quantization algorithms (AWQ, GPTQ, SmoothQuant), communication operators (Allreduce), compute-communication overlap, separated deployment (PD separation)

Preferred skills

PaddlePaddle, PyTorch, TensorFlow, vLLM, TGI, SGLang, TensorRT-LLM

Responsibilities

Optimize inference performance for Baidu's ERNIE Bot and other business LLMs; design and develop PaddlePaddle inference framework; perform deep optimization of functional modules on CPU/GPU; track and implement breakthroughs in deep learning inference technologies; optimize framework usability for developers; design and develop HPC platforms and libraries.

Seniority

Senior, hands-on IC

Sourced via baidu · Listed on CareerPlan, which tracks 877,000+ jobs from 20+ sources.