多模态算法工程师-抖音AI分身
Core
Design and implement multimodal AI solutions for Douyin, focusing on generating interactive content (video, audio, action-driven) from creator assets like live streams and short videos to enhance user engagement and business metrics.
Role type
Senior Multimodal Algorithm Engineer (Generative AI)
Builds
AI-generated video content, interactive agents, and personalized creator assets for Douyin's live streaming and short video ecosystem.
Domain
Social Media / Generative AI / Multimodal Learning
Deliverable
production ML models
Required skills
Multimodal Large Language Models (MLLM), Vision-Language Models (VLM), Diffusion models, Reinforcement Learning (RL), Supervised Fine-Tuning (SFT), Agent planning, RAG, Auto-Prompting, Data engineering, Model evaluation
Preferred skills
Experience in short video/live streaming algorithms, publications in top-tier conferences (NIPS, CVPR, ICML, etc.), competition experience
Technologies
PyTorch, TensorFlow, Hugging Face, LangChain, Vector Databases, Kubernetes
Responsibilities
Design and optimize multimodal models for content generation and classification; Build agent-based systems for video scripting and real-time interaction; Establish infrastructure for model training, inference, and evaluation; Research and adapt new generative AI techniques for business scenarios.
Seniority
Senior, hands-on IC