多模态算法工程师-文档智能
Core
Research and develop multimodal large models for document intelligence, including visual rich document parsing, key information extraction, video summarization, and visual document translation.
Role type
Research-level multimodal algorithm engineer (document intelligence)
Builds
Efficient, scalable multimodal large model architectures for document processing tasks
Domain
Computer Vision + Multimodal AI + Document Intelligence
Deliverable
production ML models
Required skills
C++, Python, computer vision algorithms, multimodal algorithms, machine learning algorithms, academic research
Preferred skills
Published papers in computer vision, patent applications
Technologies
Large language models, multimodal architectures
Responsibilities
Explore applications of multimodal understanding and generation models in document intelligence, follow up on frontier technologies, conduct deep research on key technical challenges
Seniority
Senior, research & hands-on IC
