推理性能优化专家-计算
Core
Lead inference optimization for LLMs and multimodal models, focusing on quantization, sparsity, communication protocols, and core operator efficiency to enable scalable deployment.
Role type
Senior IC inference performance optimization engineer
Builds
Standardized performance benchmarking systems, automated tuning pipelines, and high-throughput distributed inference clusters
Domain
AI/ML infrastructure, distributed systems, high-performance computing
Deliverable
production ML models
Required skills
C++/Python/Go, LLM inference optimization, mixed-precision quantization, sparse attention mechanisms, RDMA/TCP protocol optimization, AI core operator development (GEMM, convolution), Triton/MLIR compilation frameworks, hardware instruction set adaptation (CUDA/ROCm/Ascend/Cambricon), distributed system topology design, speculative decoding, MoE routing strategies
Preferred skills
PyTorch/TensorFlow kernel architecture knowledge, system-level profiling tools (Nsight, PyTorch Profiler), low-latency serialization, traffic scheduling
Technologies
INT4/INT8/FP8, Sparse Attention, RDMA, Triton, MLIR, CUDA, ROCm, Ascend, Cambricon, PyTorch Profiler, NVIDIA Nsight
Responsibilities
Implement hybrid precision quantization and sparsity techniques; optimize cross-node communication stacks and topology; develop and fuse AI core operators for hardware acceleration; design end-to-end scheduling and overlap strategies for MoE architectures
Seniority
Senior, hands-on IC