CareerPlanSign in

LLM/VLM大模型推理算子/框架优化专家 - Data AML

上海💼 Full-time🗓 2026-09-28

Core

Optimizing inference for LLMs and VLMs on heterogeneous chips to maximize large-scale compute power for the Volcano Engine Ark platform.

Role type

Senior IC LLM/VLM inference optimization engineer

Builds

High-performance inference systems for Deepseek, Kimi, GLM, and other large models on diverse chip architectures

Domain

AI Infrastructure / Heterogeneous Computing / Large Language Models

Deliverable

production ML models

Required skills

C/C++, Python, Linux, Computer Architecture, Parallel Computing, GPU/NPU Hardware Architecture, Inference Optimization Software Stacks (CUDA, CUTLASS, AscendC, BangC, HIP, FlyDSL), Heterogeneous Chip AI Model Performance Analysis, Operator Optimization

Preferred skills

LLM/VLM Architecture Knowledge, Inference Frameworks (vLLM, SGLang), Large Model Parallel Strategies, Quantization Algorithms

Technologies

CUDA, CUTLASS, AscendC, BangC, HIP, FlyDSL, vLLM, SGLang

Responsibilities

Optimize inference for LLMs/VLMs across multiple heterogeneous chips; Enhance the inference adaptation and optimization technology system for different chips; Analyze and evaluate new heterogeneous chips for large models; Research and implement cutting-edge inference acceleration and hardware-software co-optimization techniques.

Sourced via bytedance · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.