大模型推理研发专家-基础设施
Core
Build high-performance LLM inference service engines and platforms, optimizing throughput and latency while balancing cost.
Role type
Senior IC machine-learning infrastructure engineer (LLM inference)
Builds
LLM inference serving platforms and optimization toolkits
Domain
Large Language Models, System Performance Optimization, GPU Computing
Deliverable
production ML models
Required skills
C/C++, Python, Linux, LLM inference frameworks (vLLM, TensorRT-LLM, SGLang), performance profiling (Perf, eBPF, Nsight), system bottleneck analysis
Preferred skills
GPU architecture and software stack (CUDA, cuDNN), InfiniBand/RDMA network programming, framework secondary development
Technologies
vLLM, TensorRT-LLM, SGLang, Tensorflow, PyTorch, CUDA, cuDNN, InfiniBand, RDMA, eBPF, Perf, Nsight
Responsibilities
Develop LLM inference service engines and platforms; Analyze and optimize full-stack inference performance to meet SLO/SLA; Research and introduce forward-looking technical architectures like compilation optimization and model quantization.