多模态算法工程师(即梦AI方向) - 剪映CapCut
Core
Develop multimodal foundation models and application algorithms for Jiemeng AI, focusing on text-image-video understanding, temporal modeling, cross-modal alignment, and creative generation/editing.
Role type
Senior Multimodal Algorithm Engineer (Generative AI)
Builds
Multimodal creation models and AIGC products for Jiemeng and CapCut ecosystem
Domain
Generative AI, Multimodal Large Models, Computer Vision
Deliverable
production ML models
Required skills
Multimodal large models, generative foundation models, image/video understanding, VLM post-training, temporal modeling, data construction, pre-training, post-training, model evaluation, Python, PyTorch, SFT, DPO/GRPO, RLHF/RLVR, OPD, DiT, self-attention, model distillation
Preferred skills
High-impact publications in top conferences, distributed training framework development, model inference optimization, online deployment, open-source contributions, algorithm competition awards
Technologies
PyTorch, DiT, VLM, RLHF, DPO, GRPO, OPD
Responsibilities
Research and develop multimodal creation algorithms to solve complex intent understanding and generation quality issues; Build data and evaluation loops using real user feedback to optimize models; Explore multimodal creative agents with context understanding and tool collaboration capabilities
Seniority
Senior, hands-on IC