CareerPlanSign in

企业微信-大模型算法工程师-Agent后训练(北京/广州)

Chengdu, China💼 Full-time🗓 2026-09-28

Core

Designing and implementing post-training strategies for Large Language Models (LLMs) and Agents to enhance planning, reasoning, tool use, and instruction following in real-world scenarios.

Role type

Senior IC LLM/Agent post-training engineer

Builds

Production-ready LLMs and Agent systems optimized for specific business tasks

Domain

Artificial Intelligence / Large Language Models / Reinforcement Learning

Deliverable

production ML models

Required skills

LLM post-training (SFT, RLHF, DPO, PPO/GRPO), Python, PyTorch, DeepSpeed/Megatron, veRL/OpenRLHF, distributed training optimization, data synthesis and cleaning, reward modeling

Preferred skills

Experience with multi-step planning and hallucination suppression, reinforcement learning environment setup, publications in top-tier conferences (ICLR, NeurIPS, ICML, ACL, EMNLP), contributions to open-source ML projects

Sourced via tencent · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.