CareerPlanSign in

Sr. AI Inference Platform Engineer

Seattle, United States of America💼 Full-time🗓 2026-08-10 → 2026-09-28

Core

Build tooling, automation, and analysis capabilities for AI inference performance benchmarking, capacity projection, and data pipelines to inform infrastructure scaling decisions.

Role type

Senior IC infrastructure engineer (AI inference platform)

Builds

Performance benchmarking systems, capacity projection models, and data analysis pipelines

Domain

AI infrastructure, distributed systems, datacenter engineering

Deliverable

production ML models

Required skills

AI/ML inference architecture, distributed systems performance engineering, Python/Go/C++, automation engineering, GPU/accelerator architecture, statistical analysis

Preferred skills

performance benchmarking methodologies, capacity planning forecasting, GPU profiling, data visualization, ML serving frameworks (Triton, TensorRT-LLM, vLLM), CI/CD orchestration, Kubernetes, metrics/logging tools (Prometheus, Grafana, Splunk)

Technologies

Python, Go, C++, Nsight, Triton, TensorRT-LLM, vLLM, Kubernetes, Prometheus, Grafana, Splunk

Responsibilities

Design automations to evaluate AI inference performance across hardware generations; Develop tooling to surface performance trends and regressions; Build projection models for long-term capacity planning; Analyze utilization data to identify bottlenecks; Partner with infrastructure and hardware teams to deliver critical data; Create performance analysis workflows to increase team velocity; Improve accuracy and coverage of performance measurement systems

Seniority

Senior, hands-on IC

Sourced via apple · Listed on CareerPlan, which tracks 852,000+ jobs from 20+ sources.