CareerPlanSign in

腾讯云-MAAS SRE工程师

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Ensure availability and performance of MaaS inference services (gateway/scheduling/protocol adaptation/engine) under high concurrency, optimizing metrics like TTFT/TPOT/throughput and latency.

Role type

Senior IC SRE specializing in large-scale inference platform stability and optimization

Builds

High-availability MaaS inference infrastructure serving large language model workloads

Domain

Cloud Infrastructure / Large Language Model Inference

Deliverable

production ML models

Required skills

Go/Python/Shell, Kubernetes, GPU containerization, CUDA programming, RDMA/NCCL, vLLM/SGLang/TensorRT-LLM, capacity modeling, observability

Preferred skills

AI Agent for root cause analysis, heterogeneous compute orchestration, new collective communication patterns

Technologies

vLLM, SGLang, TensorRT-LLM, Triton, K8s, Docker, CUDA, RDMA, NCCL

Responsibilities

Define and achieve SLA for inference services; optimize inference engine batch scheduling and KV Cache; design multi-AZ disaster recovery architectures; build capacity models for cost optimization; construct observability and automated troubleshooting systems

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.