CareerPlanSign in

AML-火山方舟大模型推理系统工程师

杭州💼 Full-time🗓 2026-09-28

Core

Design, develop, and optimize large-scale training and inference systems for large language models, handling high-concurrency traffic and heterogeneous hardware.

Role type

Senior IC large model inference system engineer

Builds

High-performance distributed training and inference clusters for global multi-region GPU workloads

Domain

Cloud computing / Large Language Models / Distributed Systems

Deliverable

production ML models

Required skills

C/C++, Python, Linux, distributed systems architecture, GPU hardware architecture, CUDA, cuDNN, model quantization, subgraph matching, compilation optimization, task orchestration, elastic scheduling, GPU overcommitment

Preferred skills

Experience with search/ad/recommendation systems, vLLM, TensorRT-LLM, SGLang, Megatron-LM, NPU/TPU integration, PhD in computer systems

Technologies

vLLM, TensorRT-LLM, SGLang, Megatron-LM, CUDA, cuDNN, PyTorch, TensorFlow, MxNet, Linux, GPU, NPU, TPU

Responsibilities

Optimize model compute performance and tune thousand-card training clusters; Solve high-concurrency, high-reliability, and high-scalability technical challenges; Research and introduce forward-looking architecture techniques like subgraph matching and model quantization; Integrate heterogeneous hardware (GPU, NPU, TPU) into training/inference frameworks; Improve GPU cluster utilization via elastic scheduling and task orchestration; Collaborate with algorithm teams for joint optimization

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.