大模型推理服务部署框架资深工程师-智能创作(北京/上海/深圳)
Core
Design and implement architecture for multi-modal LLM/VLM/AIGC inference services, optimize inference, deploy services, and ensure high availability and low cost.
Role type
Senior IC machine-learning engineer (inference services)
Builds
High-availability, low-cost inference services for multi-modal large models
Domain
Artificial Intelligence / Large Language Models / Inference Optimization
Deliverable
production ML models
Required skills
Python, C, C++, Go, software engineering, LLM/VLM/AIGC knowledge, distributed high-concurrency system architecture, GPU/NPU hardware characteristics, Kernel development and tuning
Preferred skills
ComfyUI, VLLM, Slang, Ray, RPC
Responsibilities
Design and implement multi-modal LLM/VLM/AIGC inference service architecture; Optimize inference performance and manage service deployment; Collaborate with product teams to ensure high-standard delivery