边缘AI推理研发工程师(深圳)
Core
Design and optimize low-latency AI inference systems and KVCache platforms for real-time voice interaction and edge AI applications.
Role type
Senior IC Edge AI Inference Engineer
Builds
Serverless AI inference platforms and KVCache systems for edge cloud ecosystems
Domain
Edge computing, Large Language Models (LLM), Low-latency systems
Deliverable
production ML models
Required skills
LLM model architecture, GPU/TPU architecture, Inference acceleration, sglang, vLLM, TensorRT-LLM, KVCache optimization, Linux asynchronous IO, RDMA
Preferred skills
Industrial-scale LLM deployment experience, Heterogeneous chip adaptation, Low-latency storage optimization
Responsibilities
Optimize end-to-end inference latency through scheduling, transmission, and computation; Design and develop Serverless AI inference platforms; Build KVCache systems with offloading, compression, and sharing techniques; Optimize low-latency disk IO and network transmission.
Seniority
Senior, hands-on IC