CareerPlanSign in

Inference Engineer, AGI

Sunnyvale, California, United States💼 Full-time🗓 2026-10-05

Core

Design and optimize real-time inference systems for frontier-scale multimodal conversational AI, ensuring models run within strict latency budgets on production hardware. (via careerplan.io/jobs/10569698-inference-engineer-agi-at-amazon)

Role type

Senior Inference Engineer (Multimodal/Real-time)

Builds

Low-latency streaming runtime, offline training/RL infrastructure, and high-performance inference kernels.

Domain

AI/ML, Real-time Systems, Speech & Audio

Deliverable

production ML models

Required skills

GPU performance optimization, deep learning architectures (transformers, attention), inference efficiency techniques (quantization, speculative decoding), custom kernel development, distributed systems design, real-time latency management

Preferred skills

vLLM/TensorRT-LLM internals, CUTLASS/Triton/CUDA kernel authoring, reinforcement learning infrastructure, multi-GPU tensor parallelism, speech/audio generative models, open-source inference contributions

Technologies

vLLM, PyTorch, CUDA, CUTLASS, Triton, TensorRT-LLM, NCCL, NVLink, Nsight Compute

Responsibilities

Co-design model architectures for inference efficiency, implement and optimize critical inference path components, build real-time serving frameworks for streaming AI, develop offline systems for RLHF and evaluation, profile and eliminate performance bottlenecks in large-scale workloads

Seniority

Senior, hands-on IC