微信WeLM-大模型研发基础设施工程师-研发效能方向(深圳、上海、广州)
Core
Design and maintain large-scale production systems for model research iteration, optimizing scheduling, routing, and deployment strategies for training, inference, and evaluation workloads.
Role type
Senior IC ML Systems Infrastructure Engineer
Builds
Sandbox CPU clusters, GPU clusters, Kubernetes environments, and data systems for model research
Domain
AI/ML Infrastructure, Large Language Models (LLMs)
Deliverable
infrastructure
Required skills
Distributed systems, ML Systems, Task scheduling, Workflow systems, Containerization, Kubernetes, Storage systems, Message systems, Observability
Preferred skills
Large model training, Inference, Evaluation, Synthetic data generation, Reinforcement learning Rollout
Responsibilities
Optimize system throughput and stability for research iteration loads; Operate sandbox CPU and GPU clusters for environment isolation and task orchestration; Ensure SLA, observability, and fault recovery for production research systems; Collaborate with algorithm and inference teams to evolve system capabilities for frontier model iterations.