CareerPlanSign in

腾讯游戏-高性能算子优化工程师/专家

Hangzhou, China💼 Full-time🗓 2026-09-28

Core

Design, implement, and optimize operators for large model training and inference to maximize throughput, latency, and memory utilization.

Role type

Senior IC high-performance computing engineer (GPU operator optimization)

Builds

Optimized CUDA/CUTLASS/Triton operators for large-scale model training and inference

Domain

AI/ML infrastructure, GPU computing, distributed systems

Deliverable

production ML models

Required skills

C/C++, CUDA programming, CUTLASS/Triton frameworks, GPU architecture knowledge, distributed computing, performance profiling (Nsight, nvprof, perf), communication/computation overlap strategies

Preferred skills

Experience with compiler libraries, network hardware, system software

Technologies

CUDA, CUTLASS, Triton, Nsight, nvprof, perf

Responsibilities

Design and implement operators for large model training/inference; Perform end-to-end performance analysis and tuning; Optimize communication operators and design compute-communication overlap strategies; Collaborate with distributed training and algorithm teams; Track and integrate latest hardware and system software technologies

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.