企业微信-大模型训练框架开发工程师-AI Infra(成都/北京)
Core
Develop and optimize high-performance operators and distributed training/inference frameworks for large language models on NVIDIA GPUs and domestic chips.
Role type
Senior IC machine-learning infrastructure engineer (LLM training/inference)
Builds
Production-ready LLM training and inference frameworks optimized for GPU clusters
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
C/C++, CUDA, Triton, GPU operator development, distributed training optimization, performance profiling, NCCL, NVLink, PyTorch, TensorRT-LLM, vLLM, SGLang, Megatron-LM
Preferred skills
Experience with domestic AI accelerators, deep understanding of LLM architecture modules (Attention, MoE, KV Cache)
Technologies
NVIDIA GPU, CUDA, Triton, C++, PyTorch, TensorRT-LLM, vLLM, SGLang, Megatron-LM, NCCL, NVLink, NVSwitch, InfiniBand, RDMA, Nsight Systems, Nsight Compute, CUDA Profiler
Responsibilities
Analyze and optimize performance for LLM training/inference scenarios focusing on compute, memory, and communication bottlenecks; Develop and optimize high-performance operators (GEMM, Attention, MoE, KV Cache, communication fusion); Tune multi-card/multi-machine training/inference performance; Optimize framework implementations (PyTorch, TensorRT-LLM, vLLM, SGLang, Megatron-LM); Use profiling tools to identify and resolve performance bottlenecks.
Seniority
Senior, hands-on IC
