CareerPlanSign in

大模型推理优化工程师(上海/深圳)

Shanghai, China💼 Full-time🗓 2026-09-28

Core

Optimize online inference services for large language models (LLMs) in gaming business, focusing on low latency and high throughput across heterogeneous hardware.

Role type

Senior IC LLM inference optimization engineer

Builds

High-performance LLM inference engines and services for gaming projects

Domain

Gaming industry + Large Language Model inference

Deliverable

production ML models

Required skills

Python, C++, Linux system programming, CUDA programming, GPU performance tuning, PyTorch, vLLM, SGLang, TensorRT-LLM, quantization (FP8, INT4, AWQ, MXFP4), sparse models, operator fusion, graph compilation, profiling (Nsight), memory management, communication optimization

Preferred skills

Experience with heterogeneous hardware deployment, speculative decoding, MoE expert parallelism, long context optimization

Technologies

vLLM, SGLang, TensorRT-LLM, CUDA, Triton, Nsight, PyTorch, Linux

Responsibilities

Optimize inference latency (TTFT) and throughput; implement compression techniques like quantization and sparsity; develop and optimize inference engines and operators; explore advanced architectures for long context and high concurrency; adapt inference services to various hardware platforms; establish monitoring and profiling systems for continuous optimization.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.