Research Scientist, (Privacy-Preserving Large-Scale Model Training & Architecture Optimization)
Core
Design and optimize large-scale training architectures for diffusion-based and unified generative foundation models in privacy-sensitive production environments.
Role type
Senior IC machine-learning systems engineer (diffusion & unified models)
Builds
Next-generation generative foundation models (DiT, Rectified Flow, hybrid AR + diffusion) deployed in production
Domain
AI/ML, Large-scale distributed systems, GPU optimization, Privacy-preserving ML
Deliverable
production ML models
Required skills
Large-scale deep learning systems, Distributed training (DP/TP/PP/ZeRO/FSDP), GPU optimization (memory layout, kernel fusion), Diffusion model training, PyTorch, Fault-tolerant system design
Preferred skills
Privacy-preserving ML, CUDA kernel development, Unified multimodal models, Production GPU orchestration at scale
Technologies
PyTorch, CUDA, ZeRO, FSDP, DiT, Rectified Flow
Responsibilities
Design end-to-end training architecture for diffusion and unified models; Optimize GPU-centric performance across thousands of accelerators; Build fault-tolerant, self-healing training systems; Optimize noise schedules and memory-efficient attention mechanisms.
Seniority
Senior, hands-on IC