端侧多模态推理引擎高性能优化工程师-AML(上海/杭州/广州/深圳)
Core
Develop and optimize ByteDance's company-level on-device AI inference framework using parallel computing, architecture design, sparsity optimization, and heterogeneous scheduling to create a high-performance heterogeneous AI inference engine.
Role type
Senior IC machine-learning engineer (on-device inference optimization)
Builds
High-performance heterogeneous AI inference engine for LLM, multimodal, and AIGC models deployed in products like Douyin, Jianying, and Volcano Engine.
Domain
Mobile/Edge AI, Heterogeneous Computing, Model Optimization
Deliverable
production ML models
Required skills
C/C++, CUDA, OpenCL, Metal, TensorRT, Triton, CUTLASS, ARM NEON assembly, FlashAttention, Conv2d, GEMM, GEMV, Llama.cpp, MNN, TNN, CoreML, MoE architecture, low-bit quantization, SparseAttention
Preferred skills
Experience with Qualcomm/MTK NPU architectures, deep understanding of computer architecture, parallel computing optimization for mobile/PC/automotive platforms
Technologies
NVIDIA, Adreno, Mali, Apple GPUs, ARM/x86 CPUs, TensorRT, Triton, CUTLASS, Llama.cpp, MNN, TNN, CoreML
Responsibilities
Develop and optimize on-device AI inference frameworks via CPU/GPU/DSP/NPU parallel computing and heterogeneous scheduling; Optimize LLM, multimodal, and AIGC algorithms for on-device inference; Develop AI model and inference framework toolchains and build technical ecosystems.