CareerPlanSign in

AI Infrastructure Engineer Graduate (Algorithm Infrastructure) - 2027 Start (PhD)

San Jose, United States of America💼 Full-time🗓 2026-09-28

Core

Building and evolving next-generation inference systems for ultra-large-scale language models, vision-language models, and frontier multimodal AI systems, focusing on distributed serving, heterogeneous scheduling, and low-latency inference.

Role type

PhD-level AI Infrastructure Engineer (Algorithm Infrastructure)

Builds

High-performance inference systems for 200B+ models and complex multimodal models

Domain

AI Infrastructure / Large-Scale Model Serving

Deliverable

production ML models

Required skills

System design for high-concurrency environments, Asynchronous scheduling, Resource pooling, Load balancing, Performance optimization, CUDA programming, Triton development, Heterogeneous compute scheduling, High-concurrency load balancing, Batch formation, Kernel efficiency, Distributed inference strategies (TP, EP, DP), MoE architecture support, Emerging attention mechanisms, Multimodal fusion layers, AI-driven infrastructure development, AI Agents for optimization, Deployment pipelines, Consistency validation, Intelligent operations

Preferred skills

null

Technologies

CUDA, Triton, TP, EP, DP, MoE, Heterogeneous compute

Responsibilities

Build and evolve next-generation inference systems for large-scale online traffic, Optimize distributed inference for 200B+ models through TP, EP, DP, and related strategies, Develop high-performance kernels for frontier model architectures, Explore AI-driven infrastructure for inference systems

Seniority

PhD, Research & Engineering

Sourced via tiktok · Listed on CareerPlan, which tracks 851,000+ jobs from 20+ sources.