AML-火山方舟大模型推理系统工程师
Core
Design, develop, and optimize large-scale training and inference systems for large language models, handling high-concurrency traffic and heterogeneous hardware.
Role type
Senior IC large model inference system engineer
Builds
High-performance distributed training and inference clusters for global multi-region GPU workloads
Domain
Cloud computing / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
C/C++, Python, Linux, distributed systems architecture, GPU hardware architecture, CUDA, cuDNN, model quantization, subgraph matching, compilation optimization, task orchestration, elastic scheduling, GPU overcommitment
Preferred skills
Experience with search/ad/recommendation systems, vLLM, TensorRT-LLM, SGLang, Megatron-LM, NPU/TPU integration, PhD in computer systems
Technologies
vLLM, TensorRT-LLM, SGLang, Megatron-LM, CUDA, cuDNN, PyTorch, TensorFlow, MxNet, Linux, GPU, NPU, TPU
Responsibilities
Optimize model compute performance and tune thousand-card training clusters; Solve high-concurrency, high-reliability, and high-scalability technical challenges; Research and introduce forward-looking architecture techniques like subgraph matching and model quantization; Integrate heterogeneous hardware (GPU, NPU, TPU) into training/inference frameworks; Improve GPU cluster utilization via elastic scheduling and task orchestration; Collaborate with algorithm teams for joint optimization
Seniority
Senior, hands-on IC