VLM Research Engineer (m/f/d)
Core
Design and adapt vision-language and video models for scene understanding, temporal reasoning, and activity recognition, then build production pipelines for customer use.
Role type
Senior Research Engineer (Applied Multimodal AI)
Builds
Production-ready inference pipelines, large-scale training/evaluation systems, and robust benchmarks for video QA and instruction following.
Domain
Computer Vision, Multimodal Learning, Video Understanding
Deliverable
production ML models
Required skills
Video-centric deep learning, Large VLM training and adaptation, Multi-GPU distributed training, Python engineering, Model compression and optimization
Preferred skills
Top-tier publications in video/multimodal learning, 3D/4D scene representations, Embodied AI (sense-plan-act), Inference optimization (quantization, TensorRT)
Technologies
PyTorch, InternVL, Qwen-VL, DeepSeek-VL, GPU clusters
Responsibilities
Design and adapt VLMs for temporal reasoning and action recognition; Build and maintain large-scale training pipelines; Curate and augment video-text datasets; Develop benchmarks for video QA and temporal understanding; Refactor architectures for efficiency and deployability; Deliver production inference pipelines to product teams.
Seniority
Senior, hands-on IC