ML Algorithm Mapping and Performance Engineer, Core ML
Core
Engineer mapping ML algorithms to Cerebras architecture to optimize training/inference efficiency and characterize performance frontiers.
Role type
Senior IC ML Systems & Performance Engineer
Builds
Analytical performance models, prototype implementations, benchmarks, and performance visualization tools for Cerebras WSE and GPU baselines.
Domain
AI Hardware / Machine Learning Systems / High-Performance Computing
Deliverable
production ML models
Required skills
Computer architecture, parallel computing, ML fundamentals, analytical performance modeling, algorithmic complexity analysis, Python, C++, system profiling/debugging
Preferred skills
Roofline analysis, hardware-software co-design, CUDA/Triton/PyTorch/JAX, transformer internals, research publications
Technologies
Cerebras WSE, CUDA, Triton, PyTorch, JAX
Responsibilities
Build analytical and empirical performance models for ML algorithms; Characterize asymptotic behavior and algorithmic trade-offs; Construct Pareto frontiers across quality, latency, and cost; Develop prototype implementations and benchmarks; Analyze system behavior to identify bottlenecks; Evaluate emerging techniques like speculative decoding and MoE; Partner with research and kernel teams; Develop performance visualization tools.
Seniority
Senior, hands-on IC