CareerPlanSign in

AI Inference Platform Engineer

Chicago💼 Full-time💰 $200,000–$250,000🗓 2026-09-29 → 2026-09-30

Core

Build, operate, and optimize the systems that serve large language, vision, multimodal, and embedding models across DRW, providing the firmwide interface to modern AI models from evaluation through production use.

Role type

Senior IC AI Inference Platform Engineer

Builds

Inference runtimes, distributed systems, and production platform for LLMs and multimodal models

Domain

Financial Trading / AI Infrastructure

Deliverable

production ML models

Required skills

LLM inference optimization, NVIDIA GPU architecture expertise, inference runtime development, distributed system design, multi-tenant scheduling, performance profiling, model quality equivalence testing, CI/CD for model serving

Preferred skills

Experience with Hopper/Blackwell GPUs, knowledge of speculative decoding and quantization, Linux systems performance fundamentals

Technologies

TensorRT-LLM, vLLM, SGLang, CUDA, Nsight, DCGM, OpenTelemetry, Prometheus, Grafana

Responsibilities

Optimize LLM inference performance across modern NVIDIA GPU architectures; Build end-to-end performance profiling and observability; Design and optimize KV cache and distributed inference architectures; Own day-0 model onboarding and serving configuration; Measure and monitor quality equivalence across serving configurations; Manage the production serving lifecycle of models; Partner with SRE and platform teams to automate deployment and operation; Optimize model placement and resource allocation; Design and operate multi-tenant scheduling and isolation

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 878,000+ jobs from 20+ sources.