推理执行引擎研发工程师 - Data AML
Core
Designing and optimizing the underlying architecture of large model inference engines for high-throughput, low-latency GPU/NPU deployment.
Role type
Senior IC inference engine engineer (large models)
Builds
High-performance, low-loss inference infrastructure for large language models
Domain
AI/ML infrastructure, GPU computing, distributed systems
Deliverable
production ML models
Required skills
C/C++, CUDA, GPU architecture, operator fusion, distributed parallelism (TP/PP/SP/MoE), compilation optimization, performance profiling
Preferred skills
vLLM, TensorRT-LLM, SGLang, model parallelism strategies
Technologies
CUDA, Nsight, Profiler, TensorRT-LLM, vLLM, SGLang
Responsibilities
Optimize GPU memory access and compute pipelines to eliminate inference bottlenecks; Design and implement distributed parallel strategies for multi-card deployment; Benchmark against and innovate upon mainstream inference frameworks.
Seniority
Senior, hands-on IC