Sr. AI Inference Platform Engineer
Core
Build tooling, automation, and analysis capabilities for AI inference performance benchmarking, capacity projection, and data pipelines to inform infrastructure scaling decisions.
Role type
Senior IC infrastructure engineer (AI inference platform)
Builds
Performance benchmarking systems, capacity projection models, and data analysis pipelines
Domain
AI infrastructure, distributed systems, datacenter engineering
Deliverable
production ML models
Required skills
AI/ML inference architecture, distributed systems performance engineering, Python/Go/C++, automation engineering, GPU/accelerator architecture, statistical analysis
Preferred skills
performance benchmarking methodologies, capacity planning forecasting, GPU profiling, data visualization, ML serving frameworks (Triton, TensorRT-LLM, vLLM), CI/CD orchestration, Kubernetes, metrics/logging tools (Prometheus, Grafana, Splunk)
Technologies
Python, Go, C++, Nsight, Triton, TensorRT-LLM, vLLM, Kubernetes, Prometheus, Grafana, Splunk
Responsibilities
Design automations to evaluate AI inference performance across hardware generations; Develop tooling to surface performance trends and regressions; Build projection models for long-term capacity planning; Analyze utilization data to identify bottlenecks; Partner with infrastructure and hardware teams to deliver critical data; Create performance analysis workflows to increase team velocity; Improve accuracy and coverage of performance measurement systems
Seniority
Senior, hands-on IC