Applied Researcher: On-Device Multimodal Reasoning
Core
Design and optimize compact vision-language models (under ~10B parameters) for real-time, on-device multimodal reasoning, planning, and acting across vision, language, and 3D scene content within strict compute and memory constraints.
Role type
Senior IC Applied Researcher (Efficient Multimodal Reasoning)
Builds
Next-generation on-device AI systems that understand and act across language, vision, audio, and tools for Apple products
Domain
Consumer Electronics / Efficient Machine Learning / Computer Vision
Deliverable
production ML models
Required skills
Deep learning, Vision-Language Model (VLM) training and post-training, Model compression (distillation, pruning, quantization), Efficient inference decoding (speculative, structured), Multimodal reasoning, Python, PyTorch
Preferred skills
PhD in AI/ML, Chain-of-thought compression, RL for reasoning (GRPO, RLVR), Hardware-aware optimization, Apple Silicon expertise
Technologies
PyTorch, vLLM, SGLang, llama.cpp, MLX, CoreML, Neural Engine, GPU, ANE