企业微信-大模型算法工程师-Agent后训练(北京/广州)
Core
Designing and implementing post-training strategies for Large Language Models (LLMs) and Agents to enhance planning, reasoning, tool use, and instruction following in real-world scenarios.
Role type
Senior IC LLM/Agent post-training engineer
Builds
Production-ready LLMs and Agent systems optimized for specific business tasks
Domain
Artificial Intelligence / Large Language Models / Reinforcement Learning
Deliverable
production ML models
Required skills
LLM post-training (SFT, RLHF, DPO, PPO/GRPO), Python, PyTorch, DeepSpeed/Megatron, veRL/OpenRLHF, distributed training optimization, data synthesis and cleaning, reward modeling
Preferred skills
Experience with multi-step planning and hallucination suppression, reinforcement learning environment setup, publications in top-tier conferences (ICLR, NeurIPS, ICML, ACL, EMNLP), contributions to open-source ML projects
