Sr. AI Engineer - Inference Optimization
Core
Design, build, test, deploy, and maintain AI software components including foundation model training, LLM inference, similarity search, guardrails, and observability to improve scalability, cost efficiency, latency, and throughput in large-scale production systems.
Role type
Senior IC AI Engineer (Inference Optimization)
Builds
Scalable, high-performance AI capabilities and proprietary platforms for banking customers.
Domain
Financial services / Large-scale AI systems
Deliverable
production ML models
Required skills
LLM inference optimization, similarity search, VectorDBs, model evaluation, experimentation, governance, observability, hardware utilization optimization, latency reduction, throughput improvement, cost efficiency, Python, Go, Scala, Java, C++, C#, AWS, Google Cloud, Azure, PyTorch, Hugging Face, Nemo Guardrails
Preferred skills
Leading and mentoring engineering teams, influencing cross-functional partners, applying advanced optimization techniques to training and inference software, applying new AI research methods in production
Technologies
AWS Ultraclusters, Hugging Face, VectorDBs, Nemo Guardrails, PyTorch
Responsibilities
Partner with engineers, research scientists, technical program managers, and product managers to deliver AI-powered products; design, build, test, deploy, and maintain AI software components; invent and apply state-of-the-art LLM optimization methods; help shape the technical direction and long-term roadmap for foundational AI systems.
Seniority
Senior, hands-on IC with mentorship responsibilities