Inference Engine Development - Member of Technical Staff
Core
Building hardware-aware inference engines that adapt scheduling, memory, and execution to heterogeneous accelerators beyond standard GPU clusters.
Role type
Senior IC systems engineer (inference engine development)
Builds
Production inference engines (SGLang, vLLM) optimized for heterogeneous hardware
Domain
AI infrastructure / Heterogeneous compute systems
Deliverable
production ML models
Required skills
SGLang internals, vLLM internals, C++/CUDA systems programming, high-performance Python, parallelism strategy design, disaggregated serving architecture
Preferred skills
Experience with upstream open source contributions, working with evolving APIs
Technologies
SGLang, vLLM, C++, CUDA, Python
Responsibilities
Contribute upstream to SGLang and vLLM, improve hardware-awareness in inference engines, design bespoke parallelism and disaggregation strategies, collaborate with Accelerator Systems Software engineers
Seniority
Senior, hands-on IC