Principal Software Engineering (AI Infrastructure)
Core
Lead architecture and long-term technical strategy for TSMC's AI platform, enterprise AI gateway, multi-model inference platform, GPU infrastructure, governance, and observability.
Role type
Principal Software Engineering (AI Infrastructure)
Builds
Scalable enterprise AI infrastructure for LLM serving, GPU scheduling, model routing, model lifecycle management, RAG infrastructure, vector databases, and AI orchestration.
Domain
Semiconductor industry + Cloud-native AI infrastructure
Deliverable
production ML models
Required skills
LLM inference, GPU scheduling, multi-model AI architecture, model lifecycle management, cloud-native infrastructure, NVIDIA GPU technologies, vLLM, TensorRT-LLM, Triton Inference Server, Ray, Kubeflow, MLflow, Kubernetes, Docker, Helm, ArgoCD, service mesh, AWS, Azure, GCP, Go, Python, Java, REST APIs, microservices, SDKs
Preferred skills
Distributed-systems architecture, highly available AI services, multi-region deployment, disaster recovery, service mesh, distributed caching, event-driven architecture, global load balancing, Platform-as-a-Service, AI deployment automation
Technologies
Kubernetes, Docker, Helm, ArgoCD, vLLM, TensorRT-LLM, Triton Inference Server, Ray, Kubeflow, MLflow, AWS, Azure, GCP
Responsibilities
Lead architecture and long-term technical strategy for AI platform and GPU infrastructure; Design and build scalable enterprise AI infrastructure for LLM serving and model orchestration; Lead distributed-systems architecture for highly available AI services and multi-region deployment; Drive AI platform engineering practices involving Kubernetes and AI deployment automation; Collaborate with engineering teams in North America and Taiwan to support secure, scalable, and reliable AI services across global regions.
Seniority
Principal, hands-on IC