AI Engineer 5 (FM Hosting, LLM Inference)
Core
Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, agents, multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability.
Role type
Senior IC AI Engineer (LLM Inference & Foundation Model Optimization)
Builds
Scalable, high-performance AI infrastructure and proprietary solutions for real-time, personalized customer experiences in banking.
Domain
Financial Services / Large Language Models / AI Infrastructure
Deliverable
production ML models
Required skills
Foundation model training, LLM inference, multi-agent workflows, similarity search, model evaluation, cost-performance governance, multi-model orchestration, system optimization, hardware/software/AI expertise
Preferred skills
Leading development of AI systems with tradeoff decisions, deploying scalable AI on cloud platforms, building agentic AI systems, architecting heterogeneous AI systems, enforcing ethical AI deployment standards, dynamic inference strategies, model compression, right-sizing models and hardware
Technologies
AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, CUDA, Python, Go, Scala, Java, C++, C#
Responsibilities
Partner with cross-functional teams to deliver AI-powered products; Invent and introduce state-of-the-art foundation model optimization techniques; Contribute to the technical vision and long-term roadmap of foundational AI systems; Design and implement multi-model orchestration pipelines; Establish and lead cost-performance governance reviews; Mentor Principal and Manager-level AI engineers
Seniority
Senior, hands-on IC with mentorship responsibilities




