GPU Cluster Architect
Core
Design next-generation AI infrastructure at massive scale, shaping how tens of thousands of GPUs are interconnected, powered, cooled, and optimized across multiple data center sites.
Role type
Senior IC GPU Cluster Architect
Builds
Large-scale AI and machine-learning workloads, including large language model training and inference
Domain
AI Infrastructure / High-Performance Computing
Deliverable
infrastructure
Required skills
GPU cluster architecture, high-performance networking (InfiniBand, RoCE), systems architecture, hardware reliability, workload modeling, automation scripting (Python/Go), cross-functional collaboration
Preferred skills
None explicitly stated
Technologies
NVIDIA, AMD, InfiniBand, RoCEv2, Ethernet
Responsibilities
Architect scalable GPU cluster topologies; Define infrastructure architectures for multi-site AI workloads; Model workload requirements for LLM training/inference; Design high-throughput, low-latency networking architectures; Partner with storage teams to optimize for training datasets; Analyze telemetry to identify design issues; Collaborate with SRE and data center teams to deploy architectures; Contribute to automation and telemetry initiatives; Balance scalability, performance, and operational complexity.
Seniority
Senior, hands-on IC