多模态视频理解大模型算法工程师-视频与边缘
Core
Develop and optimize frontier multimodal video understanding large models, focusing on long-video comprehension, token compression, script restoration, and real-time streaming analysis for global markets.
Role type
Senior IC multimodal large model algorithm engineer (video & edge)
Builds
Production multimodal large models (text, image, audio) and video editing agents
Domain
AI / Computer Vision / NLP / Multimodal
Deliverable
production ML models
Required skills
Python, C++, Large model training, Reinforcement Learning (RL), Video captioning, Video grounding, Video summarization, Video question answering, Action recognition
Preferred skills
Video editing, Video stylization, LongCoT, RL-based autonomous exploration, Top-tier conference publications (CVPR/ICLR/ICCV/PAMI/ACL)
Technologies
Qwen-VL, InternVL, InternVideo
Responsibilities
Innovate algorithms for long-video understanding and token compression; Optimize post-training for multimodal models using user behavior data; Research video editing agents and data synthesis; Solve industrial efficiency challenges in live streaming scenarios; Publish technical innovations via papers or patents.
