CareerPlanSign in

Staff Machine Learning Engineer, Siri Runtime Systems and Interaction

Zurich, Switzerland💼 Full-time🗓 2026-09-01 → 2026-09-28

Core

Lead the development of audio and video generation capabilities for realistic, expressive synthetic speech and visual representations to enhance human-computer interaction.

Role type

Staff Machine Learning Engineer (Generative AI, Multimodal)

Builds

End-to-end systems for generating synthetic audio (speech, acoustics) and video for agent interactions

Domain

Generative AI, Speech Synthesis, Computer Vision

Deliverable

production ML models

Required skills

Generative audio/video architectures (diffusion, autoregressive, GANs, VAEs), Model evaluation and perceptual quality assessment, Technical architecture and tradeoffs, Mentoring engineers, Research-to-production leadership

Preferred skills

Speech synthesis (TTS), Voice conversion, Generative video/animation (facial animation, lip-sync), Multimodal modeling, Low-latency model deployment, Publications in top-tier venues

Technologies

Python, PyTorch, TensorFlow, Distributed training infrastructure

Sourced via apple · Listed on CareerPlan, which tracks 852,000+ jobs from 20+ sources.