AI Research Engineer (Kernel & Inference Optimization)
Core
Develop and optimize model-serving architectures for advanced AI systems, focusing on latency, throughput, and memory efficiency across mobile and edge devices.
Role type
Senior IC AI Research Engineer (Kernel & Inference Optimization)
Builds
High-performance inference pipelines and custom GPU kernels for text, image, audio, diffusion, and vision transformer models.
Domain
AI Systems Engineering / High-Performance Computing / Mobile & Edge AI
Deliverable
production ML models
Required skills
Metal Shading Language (MSL), low-level kernel optimization, inference optimization (pruning, quantization, Flash Attention, KV caching, speculative decoding), distributed inference (tensor/pipeline/expert parallelism), GPU programming for mobile, diffusion models, Vision Transformers
Preferred skills
PhD in NLP/ML, publications at leading conferences, experience with large-scale GPU clusters
Technologies
Metal Shading Language, Flash Attention, EAGLE, tensor parallelism, pipeline parallelism, expert parallelism
Responsibilities
Design and deploy advanced model-serving architectures; Develop inference pipelines for resource-constrained environments; Build and execute controlled inference benchmarks; Identify computational bottlenecks and implement system-level optimizations; Develop custom GPU kernels and compute shaders for mobile hardware; Design and optimize distributed inference systems; Monitor production performance and refine optimization strategies.
Seniority
Senior, hands-on IC