Multimodal AI Researcher
Core
Developing foundation models for generative AI and multimodal systems that integrate real-time sensor data (video, audio) with other modalities like text to power human-centric solutions.
Role type
Applied research IC multimodal AI researcher
Builds
Multimodal generative AI models, agents, and real-time perception systems for Apple products (iPhone, Apple Vision Pro)
Domain
Consumer electronics, generative AI, computer vision, multimodal learning
Deliverable
production ML models
Required skills
foundation model development, multimodal perception systems, LLMs, VLMs, Python, PyTorch
Preferred skills
training/tuning foundation models and multimodal LLMs, training generative architectures (diffusion, RL, flow matching), real-time/streaming multimodal models, speech understanding and generation, reinforcement learning for post-training
Technologies
PyTorch, LLMs, VLMs, diffusion models, reinforcement learning
Responsibilities
Conduct algorithm research and development for multimodal generative AI and agents; drive data requirements, validation strategies, and key performance indicators; collaborate with experts in AI, ML, software, and hardware to tackle fundamental challenges.
Seniority
Mid-level, hands-on IC