CareerPlanSign in

Member of Technical Staff, Local Inference & Kernels

New York City💼 Full-time💰 $250,000–$250,000🗓 2026-09-18 → 2026-09-30

Core

Inference and performance engineer optimizing local AI models for Apple Silicon to ensure speed, efficiency, and responsiveness on Mac devices.

Role type

Senior IC machine-learning inference engineer (GPU kernels)

Builds

Local inference stack for consumer Macs (Twin product)

Domain

Consumer hardware, AI/ML inference, Apple Silicon

Deliverable

production ML models

Required skills

C++, GPU programming (Metal, CUDA, Triton), GPU architecture (memory hierarchies, bandwidth, SIMD), kernel optimization (tiling, fusion), quantization and mixed precision, profiling, matrix multiplication/attention optimization

Preferred skills

MLX internals, Metal Shading Language, llama.cpp contributions, quantized matrix multiplication, fused attention, KV-cache management, vision-language models, streaming inference

Technologies

MLX, Metal, Apple Silicon, macOS

Responsibilities

Profile latency, throughput, memory, and energy use across workloads; build optimized kernels for critical operations; evaluate quantization strategies; improve runtime allocation and scheduling; co-design model architecture with researchers; maintain performance benchmarks and numerical checks

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 877,000+ jobs from 20+ sources.