Senior Machine Learning Engineer (Large Systems)
Core
Develop and optimize AI models for Graphcore's specialized hardware, scaling to thousands of accelerators to advance large-scale AI systems.
Role type
Senior IC machine learning engineer (large-scale AI systems)
Builds
Production AI models and reference applications optimized for Graphcore's AI compute stack
Domain
Semiconductor hardware + AI compute infrastructure
Deliverable
production ML models
Required skills
Deep learning frameworks (PyTorch/JAX), Python/C++ development, model training/optimization/evaluation, experimental design, performance bottleneck analysis
Preferred skills
MLOps on Kubernetes, LLM production systems, low-precision arithmetic, C++/Triton/CUDA kernel development, distributed training/inference (64+ accelerators), HPC networking (Infiniband/NVLink/RoCE), open-source contributions
Technologies
PyTorch, JAX, Python, C++, Triton, CUDA, Kubernetes, Infiniband, NVLink, RoCE
Responsibilities
Implement and optimize ML models for performance and accuracy across thousands of accelerators; test and evaluate software releases; benchmark models to identify bottlenecks; design and execute experiments on novel AI methods; collaborate with cross-functional teams on next-gen AI hardware; engage with the AI community
Seniority
Senior, hands-on IC