CareerPlanSign in

边缘AI推理研发工程师(深圳)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Design and optimize low-latency AI inference systems and KVCache platforms for real-time voice interaction and edge AI applications.

Role type

Senior IC Edge AI Inference Engineer

Builds

Serverless AI inference platforms and KVCache systems for edge cloud ecosystems

Domain

Edge computing, Large Language Models (LLM), Low-latency systems

Deliverable

production ML models

Required skills

LLM model architecture, GPU/TPU architecture, Inference acceleration, sglang, vLLM, TensorRT-LLM, KVCache optimization, Linux asynchronous IO, RDMA

Preferred skills

Industrial-scale LLM deployment experience, Heterogeneous chip adaptation, Low-latency storage optimization

Responsibilities

Optimize end-to-end inference latency through scheduling, transmission, and computation; Design and develop Serverless AI inference platforms; Build KVCache systems with offloading, compression, and sharing techniques; Optimize low-latency disk IO and network transmission.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.