AI Infra高级工程师
Core
End-to-end performance optimization, algorithm innovation, and ecosystem deployment for large models (LLM/multimodal/reinforcement learning) on Ascend chips, covering training to inference.
Role type
Senior IC AI Infrastructure Engineer (Ascend/Hardware-Software Co-design)
Builds
High-throughput, low-latency training and inference solutions for LLMs on Ascend hardware
Domain
AI Infrastructure / Semiconductor Hardware / Large Language Models
Deliverable
production ML models
Required skills
LLM training and inference optimization, distributed training, low-precision training (FP8/FP4), long-sequence inference, hardware-software co-design, operator optimization (Matmul/Attention), C/C++, Python, computer architecture knowledge
Preferred skills
Container application development, distributed job performance optimization, collective communication optimization, open-source community planning and operations, academic/research institution collaboration
Technologies
Ascend, vLLM, SGLang, Triton, LLaMA-Factory, FlashAttention, FP8, FP4
Responsibilities
Optimize pre-training, fine-tuning, and inference workflows for mainstream LLMs on Ascend; design and implement key operators for peak compute utilization; build native support for open-source frameworks and drive commercial adoption
Seniority
Senior, hands-on IC