昇腾计算高级技术专家
Core
Design and iterate AI software stack architecture including compilers, runtimes, operator libraries, and heterogeneous chip adaptation for large model training and inference.
Role type
Senior IC AI software stack architect (compiler/runtime/performance)
Builds
AI compilers, training/inference runtimes, high-performance operator libraries, and performance tuning tools for heterogeneous chips
Domain
AI infrastructure, high-performance computing, heterogeneous computing
Deliverable
production ML models
Required skills
Linux system programming, multi-threaded high concurrency, memory optimization, assembly and instruction-level tuning, LLVM/MLIR, IR rewriting, backend code generation, CUDA/NPU kernel optimization, GEMM optimization, vectorization, memory access optimization, low-precision acceleration, collective communication, tensor/pipeline parallelism, communication compression, topology optimization, profiling, PyTorch, TensorFlow
Preferred skills
Heterogeneous AI chip software-hardware co-development, model compilation and deployment
Technologies
MLIR, LLVM, TVM, Triton, CUDA, PyTorch, TensorFlow
Responsibilities
Design AI compiler and runtime architecture, lead core module R&D for graph optimization and operator fusion, optimize full-link performance for multi-card distributed and large model inference, track and implement frontier technologies, collaborate on software-hardware co-adaptation and model deployment
Seniority
Senior, hands-on IC