Inference Engineer, AGI
Core
Design and optimize real-time inference systems for frontier-scale multimodal conversational AI, ensuring models run within strict latency budgets on production hardware. (via careerplan.io/jobs/10569698-inference-engineer-agi-at-amazon)
Role type
Senior Inference Engineer (Multimodal/Real-time)
Builds
Low-latency streaming runtime, offline training/RL infrastructure, and high-performance inference kernels.
Domain
AI/ML, Real-time Systems, Speech & Audio
Deliverable
production ML models
Required skills
GPU performance optimization, deep learning architectures (transformers, attention), inference efficiency techniques (quantization, speculative decoding), custom kernel development, distributed systems design, real-time latency management
Preferred skills
vLLM/TensorRT-LLM internals, CUTLASS/Triton/CUDA kernel authoring, reinforcement learning infrastructure, multi-GPU tensor parallelism, speech/audio generative models, open-source inference contributions
Technologies
vLLM, PyTorch, CUDA, CUTLASS, Triton, TensorRT-LLM, NCCL, NVLink, Nsight Compute
Responsibilities
Co-design model architectures for inference efficiency, implement and optimize critical inference path components, build real-time serving frameworks for streaming AI, develop offline systems for RLHF and evaluation, profile and eliminate performance bottlenecks in large-scale workloads
Seniority
Senior, hands-on IC