内容理解与生成算法工程师-Data
Core
Redesigning and upgrading traditional content understanding capabilities (ASR, OCR, face detection, voiceprint recognition, music understanding) using large model paradigms, while exploring the limits of multi-modal/omni-modal large models for business and general domains.
Role type
Senior IC multi-modal large model algorithm engineer
Builds
High-availability, middleware-based audio/video understanding and generation infrastructure for Douyin, e-commerce, live streaming, and intelligent customer service.
Domain
AI / Large Language Models / Multi-modal / Audio-Video
Deliverable
production ML models
Required skills
Machine learning fundamentals, Large Language Models (LLM), Reinforcement Learning (RL), Omni-modal models, Diffusion models, Flow-Matching, Distributed engineering frameworks, ms-Swift, VeRL
Preferred skills
Publications in top-tier conferences (NeurIPS, ICML, CVPR, etc.), Top-tier algorithm competition awards
Responsibilities
Train and maintain general multi-modal large models for specific business scenarios; Iterate data engineering, training/inference frameworks, and evaluation metrics; Drive joint modeling of understanding and generation models; Research and implement new multi-modal technologies.
Seniority
Senior, hands-on IC