大模型推理后台开发工程师(深圳/北京/上海/杭州)
Core
Design and evolve an industry-leading online inference platform for large models, supporting billions of daily calls with high performance, availability, and scalability.
Role type
Senior IC backend engineer (large model inference platform)
Builds
High-performance, high-availability inference services and standardized frameworks for model deployment
Domain
AI / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
Golang, C++, Python, Linux distributed systems, dynamic scheduling, resource management, service orchestration, long context management, system observability
Preferred skills
vLLM integration and tuning, multi-modal streaming service design, large-scale GPU cluster governance
Technologies
Golang, C++, Python, vLLM, Linux
Responsibilities
Design high-performance inference service architecture optimizing dynamic scheduling and resource management; Develop standardized inference service frameworks and toolchains for full-link deployment; Build high-availability architecture and observability systems with fault tolerance and rate limiting.
Seniority
Senior, hands-on IC