981 open positions
Engineering Manager, GPU Reliability, Accelerators↗
GoogleSunnyvale, CA, US
$207k–$207k7d
Lead the GPU Reliability Systems Operations and Tooling software organization to ensure the reliability of Google's massive GPU supercomputer systems used for AI advancements.
Senior Performance Engineer, Efficiency Red Team
AmazonSeattle, Washington, United States
7d
Hunting for hidden waste across the world's largest compute infrastructure to turn findings into freed capacity and cost savings.
Software Engineer- AI/ML, Amazon Neuron Training
AmazonCupertino, California, United States
$165k–$224k7d
Building distributed training infrastructure, parallelism techniques, and high-performance kernels for large-scale pretraining, post-training, and reinforcement learning workloads on AWS Trainium custom ML accelerators.
Sr. Software Engineer- AI/ML, Amazon Neuron Training
AmazonCupertino, California, United States
$193k–$262k7d
Lead the effort to build distributed training and post-training support for PyTorch and JAX on AWS Trainium accelerators, enabling large-scale training, post-training, and reinforcement learning workloads.
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
AmazonCupertino, California, United States
7d
Architect and implement business-critical features for distributed training on AWS Trainium, optimizing throughput and convergence for frontier-scale models.
Software Engineer Uber Technical Lead, Infrastructure, FMOD↗
GoogleSunnyvale, CA, US
$262k–$262k8d
Lead capacity planning, workload prediction, and infrastructure architecture for Google's AI and Computing Infrastructure (ACI) to optimize resource distribution and efficiency.
Software Development Engineer, Alexa Excellence, Alexa LLM Inference, Capacity, & Efficiency
AmazonBellevue, Washington, United States
$144k–$144k8d
Design, develop, and maintain large-scale distributed systems and infrastructure services for LLM inference, focusing on GPU fleet management, optimization, and cost efficiency for Alexa.
Software Engineer Uber Technical Lead, Infrastructure, FMOD↗
GoogleKirkland, WA, US
$262k–$262k8d
Lead advanced VM shape-based capacity planning and workload prediction to optimize regional and zonal resource distribution for Google's AI and Computing Infrastructure (ACI) customers.
SDE II, ML Infra Services, Annapurna Labs
AmazonSeattle, Washington, United States
$144k–$194k10d
Lead the design and implementation of an ML infrastructure platform to run, optimize, and analyze machine learning workloads on AWS ML accelerators (Inferentia, Trainium, Neuron).
Senior Performance Engineer, Efficiency Red Team↗
AmazonVancouver, BC, CA
11d
Senior performance engineer hunting for hidden waste across Amazon's largest compute infrastructure to unlock capacity and save millions in costs.
Principal Technical Account Manager, ES - NAMER - US-Frontier AI
AmazonSan Francisco, California, United States
11d
Strategic partner guiding Frontier AI research customers in designing, building, and operating large-scale AI/ML training and inference solutions on AWS.
← Select a job to preview