CareerPlanSign in

企业微信-大模型训练框架开发工程师-AI Infra(成都/北京)

Guangzhou, China💼 Full-time🗓 2026-09-28

Core

Develop and optimize high-performance operators and distributed training/inference frameworks for large language models on NVIDIA GPUs and domestic chips.

Role type

Senior IC machine-learning infrastructure engineer (LLM training/inference)

Builds

Production-ready LLM training and inference frameworks optimized for GPU clusters

Domain

Artificial Intelligence / Large Language Models / High-Performance Computing

Deliverable

production ML models

Required skills

C/C++, CUDA, Triton, GPU operator development, distributed training optimization, performance profiling, NCCL, NVLink, PyTorch, TensorRT-LLM, vLLM, SGLang, Megatron-LM

Preferred skills

Experience with domestic AI accelerators, deep understanding of LLM architecture modules (Attention, MoE, KV Cache)

Technologies

NVIDIA GPU, CUDA, Triton, C++, PyTorch, TensorRT-LLM, vLLM, SGLang, Megatron-LM, NCCL, NVLink, NVSwitch, InfiniBand, RDMA, Nsight Systems, Nsight Compute, CUDA Profiler

Responsibilities

Analyze and optimize performance for LLM training/inference scenarios focusing on compute, memory, and communication bottlenecks; Develop and optimize high-performance operators (GEMM, Attention, MoE, KV Cache, communication fusion); Tune multi-card/multi-machine training/inference performance; Optimize framework implementations (PyTorch, TensorRT-LLM, vLLM, SGLang, Megatron-LM); Use profiling tools to identify and resolve performance bottlenecks.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.