Member of Technical Staff, Local Inference & Kernels
Core
Inference and performance engineer optimizing local AI models for Apple Silicon to ensure speed, efficiency, and responsiveness on Mac devices.
Role type
Senior IC machine-learning inference engineer (GPU kernels)
Builds
Local inference stack for consumer Macs (Twin product)
Domain
Consumer hardware, AI/ML inference, Apple Silicon
Deliverable
production ML models
Required skills
C++, GPU programming (Metal, CUDA, Triton), GPU architecture (memory hierarchies, bandwidth, SIMD), kernel optimization (tiling, fusion), quantization and mixed precision, profiling, matrix multiplication/attention optimization
Preferred skills
MLX internals, Metal Shading Language, llama.cpp contributions, quantized matrix multiplication, fused attention, KV-cache management, vision-language models, streaming inference
Technologies
MLX, Metal, Apple Silicon, macOS
Responsibilities
Profile latency, throughput, memory, and energy use across workloads; build optimized kernels for critical operations; evaluate quantization strategies; improve runtime allocation and scheduling; co-design model architecture with researchers; maintain performance benchmarks and numerical checks
Seniority
Senior, hands-on IC