GPU Systems Engineer
Core
Design, deploy, and operate large-scale distributed GPU clusters for compute, storage, and AI workloads in a quantitative trading environment.
Role type
Senior GPU Systems Engineer (Infrastructure)
Builds
Distributed GPU clusters, automation tooling, and high-performance computing infrastructure for trading and research teams.
Domain
Quantitative Trading / High-Performance Computing / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, GPU workload troubleshooting, Python automation, CUDA/C++ debugging, GPUDirect RDMA, configuration management (Salt/Ansible/Puppet/Chef), network topology design, performance profiling.
Preferred skills
NVIDIA stack (NCCL, NVLink), vendor qualification, cross-layer hardware/OS/network diagnostics.
Responsibilities
Design and scale distributed GPU clusters from hardware selection to production operation; track down performance bottlenecks across compute, storage, and network; partner with researchers to profile GPU workloads and optimize speedups; build automation for provisioning, monitoring, and self-healing of thousands of nodes; qualify new hardware and software generations; root-cause complex issues with vendors.
Seniority
Senior, hands-on IC