腾讯云 EdgeOne-推理平台高级工程师
Core
Design and optimize large model inference services to improve latency, throughput, and GPU utilization.
Role type
Senior IC LLM inference platform engineer
Builds
High-performance LLM inference serving platform
Domain
AI/LLM inference systems and GPU computing
Deliverable
production ML models
Required skills
C/C++/Rust/Python, Linux system programming, LLM inference mechanisms (Prefill/Decode/Batching/KV Cache), distributed inference strategies (TP/PP/EP), inference engine optimization (TensorRT-LLM/vLLM/SGLang/Triton), CUDA/GPU performance tuning, performance analysis and bottleneck identification
Preferred skills
Experience optimizing large model inference systems
Technologies
TensorRT-LLM, vLLM, SGLang, Triton Inference Server, ORCA, CUDA
Responsibilities
Optimize inference scheduling and performance metrics (latency, throughput, memory utilization); Implement core capabilities for request scheduling, dynamic batching, and KV cache management; Optimize distributed inference strategies based on model structure and hardware; Enhance inference engine stability and performance around model loading and operator execution; Integrate and evolve frontier technologies in LLM serving and GPU optimization