CareerPlanSign in

AI Research Engineer (Kernel & Inference Optimization)

Israel🌐 Remote💼 Full-time🗓 2026-09-30 → 2026-10-01

Core

Develop and optimize model-serving architectures for advanced AI systems, focusing on latency, throughput, and memory efficiency across mobile and edge devices.

Role type

Senior IC AI Research Engineer (Kernel & Inference Optimization)

Builds

High-performance inference pipelines and custom GPU kernels for text, image, audio, diffusion, and vision transformer models.

Domain

AI Systems Engineering / High-Performance Computing / Mobile & Edge AI

Deliverable

production ML models

Required skills

Metal Shading Language (MSL), low-level kernel optimization, inference optimization (pruning, quantization, Flash Attention, KV caching, speculative decoding), distributed inference (tensor/pipeline/expert parallelism), GPU programming for mobile, diffusion models, Vision Transformers

Preferred skills

PhD in NLP/ML, publications at leading conferences, experience with large-scale GPU clusters

Technologies

Metal Shading Language, Flash Attention, EAGLE, tensor parallelism, pipeline parallelism, expert parallelism

Responsibilities

Design and deploy advanced model-serving architectures; Develop inference pipelines for resource-constrained environments; Build and execute controlled inference benchmarks; Identify computational bottlenecks and implement system-level optimizations; Develop custom GPU kernels and compute shaders for mobile hardware; Design and optimize distributed inference systems; Monitor production performance and refine optimization strategies.

Seniority

Senior, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 902,000+ jobs from 20+ sources.