大模型推理优化工程师(上海/深圳)
Core
Optimize online inference services for large language models (LLMs) in gaming business, focusing on low latency and high throughput across heterogeneous hardware.
Role type
Senior IC LLM inference optimization engineer
Builds
High-performance LLM inference engines and services for gaming projects
Domain
Gaming industry + Large Language Model inference
Deliverable
production ML models
Required skills
Python, C++, Linux system programming, CUDA programming, GPU performance tuning, PyTorch, vLLM, SGLang, TensorRT-LLM, quantization (FP8, INT4, AWQ, MXFP4), sparse models, operator fusion, graph compilation, profiling (Nsight), memory management, communication optimization
Preferred skills
Experience with heterogeneous hardware deployment, speculative decoding, MoE expert parallelism, long context optimization
Technologies
vLLM, SGLang, TensorRT-LLM, CUDA, Triton, Nsight, PyTorch, Linux
Responsibilities
Optimize inference latency (TTFT) and throughput; implement compression techniques like quantization and sparsity; develop and optimize inference engines and operators; explore advanced architectures for long context and high concurrency; adapt inference services to various hardware platforms; establish monitoring and profiling systems for continuous optimization.
Seniority
Senior, hands-on IC