CareerPlanSign in

大模型推理研发专家-基础设施

杭州💼 Full-time🗓 2026-09-28

Core

Build high-performance LLM inference service engines and platforms, optimizing throughput and latency while balancing cost.

Role type

Senior IC machine-learning infrastructure engineer (LLM inference)

Builds

LLM inference serving platforms and optimization toolkits

Domain

Large Language Models, System Performance Optimization, GPU Computing

Deliverable

production ML models

Required skills

C/C++, Python, Linux, LLM inference frameworks (vLLM, TensorRT-LLM, SGLang), performance profiling (Perf, eBPF, Nsight), system bottleneck analysis

Preferred skills

GPU architecture and software stack (CUDA, cuDNN), InfiniBand/RDMA network programming, framework secondary development

Technologies

vLLM, TensorRT-LLM, SGLang, Tensorflow, PyTorch, CUDA, cuDNN, InfiniBand, RDMA, eBPF, Perf, Nsight

Responsibilities

Develop LLM inference service engines and platforms; Analyze and optimize full-stack inference performance to meet SLO/SLA; Research and introduce forward-looking technical architectures like compilation optimization and model quantization.

Sourced via bytedance · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.