腾讯游戏-高性能算子优化工程师/专家
Core
Design, implement, and optimize operators for large model training and inference to maximize throughput, latency, and memory utilization.
Role type
Senior IC high-performance computing engineer (GPU operator optimization)
Builds
Optimized CUDA/CUTLASS/Triton operators for large-scale model training and inference
Domain
AI/ML infrastructure, GPU computing, distributed systems
Deliverable
production ML models
Required skills
C/C++, CUDA programming, CUTLASS/Triton frameworks, GPU architecture knowledge, distributed computing, performance profiling (Nsight, nvprof, perf), communication/computation overlap strategies
Preferred skills
Experience with compiler libraries, network hardware, system software
Technologies
CUDA, CUTLASS, Triton, Nsight, nvprof, perf
Responsibilities
Design and implement operators for large model training/inference; Perform end-to-end performance analysis and tuning; Optimize communication operators and design compute-communication overlap strategies; Collaborate with distributed training and algorithm teams; Track and integrate latest hardware and system software technologies
Seniority
Senior, hands-on IC