ML Runtime and Kernel Engineer - Core ML
Core
Develop novel machine learning algorithms and efficient execution on Cerebras Wafer-Scale Engine systems, bridging research ideas with high-performance runtime and kernel implementations.
Role type
Senior IC machine learning systems engineer (runtime & kernels)
Builds
High-performance ML runtimes, compilers, and low-level kernels for LLM training and inference on Cerebras hardware
Domain
AI hardware systems, machine learning infrastructure, high-performance computing
Deliverable
production ML models
Required skills
C++, Python, parallel programming, memory management, concurrency, performance optimization, system profiling and debugging, ML frameworks (PyTorch/JAX)
Preferred skills
CUDA, Triton, low-level assembly, compiler internals, distributed runtimes, HPC systems, LLM training/inference (attention, KV-cache), open-source contributions
Technologies
Cerebras Wafer-Scale Engine, PyTorch, JAX, CUDA, Triton
Responsibilities
Design and implement runtime components and high-performance kernels; Translate research prototypes into efficient implementations; Profile and debug performance across framework, compiler, runtime, and kernel layers; Optimize computation, memory movement, and concurrency; Develop benchmarks and automated tests; Collaborate with researchers and engineers on design alternatives; Contribute to software architecture and roadmap decisions
Seniority
Senior, hands-on IC
