CareerPlanSign in

推理GPU性能优化专家 - Seed Model

北京💼 Full-time🗓 2026-09-28

Core

Develop and optimize company-level large model inference frameworks and build high-performance LLM inference engines using GPU and CUDA optimizations.

Role type

Senior IC machine-learning systems engineer (LLM inference)

Builds

High-performance LLM inference engines for company-wide large models

Domain

AI / Large Language Models / GPU Computing

Deliverable

production ML models

Required skills

C/C++, Python, GPU high-performance computing, CUDA, computer architecture, parallel computing, memory optimization, low-bit computation, deep learning frameworks, neural network operators

Preferred skills

TensorRT-LLM, ORCA, vLLM, algorithm-system joint optimization

Responsibilities

Develop and optimize company-level large model inference frameworks; Build high-performance LLM inference engines via GPU/CUDA optimization; Research and introduce forward-looking ML system technologies; Collaborate with algorithm teams for joint optimization

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.