CareerPlanSign in

异构加速框架工程师(深圳/北京/上海)

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Design and optimize GPU/AI chip performance for AI inference, collaborating with algorithm teams to build high-performance operators and framework layers.

Role type

Senior IC machine-learning framework engineer (GPU/AI chip optimization)

Builds

High-performance inference frameworks and optimized AI operators

Domain

AI Infrastructure / High-Performance Computing

Deliverable

production ML models

Required skills

C/C++, Python, CUDA, Triton, Ascend C, Cublas, Cutlass, CK, Torch-Compile, parallel computing, memory optimization, communication optimization

Preferred skills

Dynamic computation graph compilation, MOE models, KV Cache optimization, dynamic batching, custom Attention operators

Technologies

CUDA, Triton, Ascend C, Cublas, Cutlass, CK, Torch-Compile

Responsibilities

Co-design GPU/AI chip performance optimizations with algorithm teams; Innovate and optimize core modules in ML frameworks; Design and implement high-performance operators and feature enablement components; Explore frontier technologies like MOE and dynamic graph compilation.

Sourced via tencent · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.