Principal Engineer - AI Networking
Core
Design, develop, and optimize RDMA-based software components and services for large-scale AI infrastructure, focusing on collective communication frameworks and transport layers.
Role type
Principal Engineer - AI Networking (Systems Programming)
Builds
Collective communication frameworks, transport layers, and communication libraries for distributed AI workloads.
Domain
AI Infrastructure / High-Performance Networking
Deliverable
production ML models | infrastructure
Required skills
Systems programming, RDMA technologies (RoCEv2, InfiniBand), C/C++, Linux systems programming, performance tuning, distributed systems concepts, networking fundamentals, debugging complex systems.
Preferred skills
Collective communication frameworks (NCCL, RCCL, MPI, UCX), AI/ML infrastructure support, GPUDirect RDMA, congestion management, Kubernetes, distributed training frameworks (PyTorch, DeepSpeed, TensorFlow), observability tools.
Technologies
InfiniBand, RoCEv2, C, C++, Linux, Kubernetes, PyTorch, TensorFlow, NCCL, MPI, UCX, GPUDirect
Responsibilities
Design and implement scalable distributed systems for AI training/inference, develop congestion management and failover capabilities, analyze and enhance communication performance across stacks, resolve complex networking and reliability issues, contribute to architectural design discussions.
Seniority
Principal, hands-on IC with strategic input
