腾讯云-MAAS SRE工程师
Core
Ensure availability and performance of MaaS inference services (gateway/scheduling/protocol adaptation/engine) under high concurrency, optimizing metrics like TTFT/TPOT/throughput and latency.
Role type
Senior IC SRE specializing in large-scale inference platform stability and optimization
Builds
High-availability MaaS inference infrastructure serving large language model workloads
Domain
Cloud Infrastructure / Large Language Model Inference
Deliverable
production ML models
Required skills
Go/Python/Shell, Kubernetes, GPU containerization, CUDA programming, RDMA/NCCL, vLLM/SGLang/TensorRT-LLM, capacity modeling, observability
Preferred skills
AI Agent for root cause analysis, heterogeneous compute orchestration, new collective communication patterns
Technologies
vLLM, SGLang, TensorRT-LLM, Triton, K8s, Docker, CUDA, RDMA, NCCL
Responsibilities
Define and achieve SLA for inference services; optimize inference engine batch scheduling and KV Cache; design multi-AZ disaster recovery architectures; build capacity models for cost optimization; construct observability and automated troubleshooting systems
Seniority
Senior, hands-on IC