Member of Technical Staff, Performance and Scale
Core
Design and implement foundational distributed systems layers to enable vLLM to serve AI models across thousands of accelerators with minimal latency and maximum reliability.
Role type
Senior IC infrastructure engineer (distributed systems)
Builds
Distributed inference infrastructure for vLLM
Domain
AI inference, high-performance computing, distributed systems
Deliverable
production ML models
Required skills
Systems programming in Rust/Go/C++, designing high-performance distributed systems, network protocols, high-performance I/O, debugging complex distributed systems
Preferred skills
ML serving infrastructure, disaggregated inference architecture, GPU programming models, GPU memory hierarchies, GPU interconnects (NVLink, InfiniBand, RoCE), improving system reliability at scale
Technologies
vLLM, Rust, Go, C++, NVLink, InfiniBand, RoCE
Responsibilities
Design foundational layers for distributed inference, implement systems serving models across thousands of accelerators, optimize for minimal latency and maximum reliability
Seniority
Senior, hands-on IC