AI/Ml Engineer
Core
Design and ship production systems for LLM features, agent workflows, retrieval, and memory to support real-time interactions.
Role type
Senior IC AI/ML Engineer (LLM & Agents)
Builds
Production LLM systems with persistent memory and real-time capabilities
Domain
Generative AI, Large Language Models, Real-time Systems
Deliverable
production ML models | product features
Required skills
Python (asyncio, FastAPI), Deep Learning (Transformers, PyTorch), LLM Fine-tuning (SFT, LoRA/QLoRA), RAG, Agent Frameworks (LangChain, AutoGen), Kubernetes, Cloud Infrastructure, Observability, Algorithmic Complexity
Preferred skills
Real-time voice/speech experience, LLM serving tools (vLLM, TensorRT-LLM), Distributed training, Open-source contributions
Technologies
PyTorch, TensorFlow, Rust, Go, C++, vLLM, TensorRT-LLM, Triton, LangChain, AutoGen, CrewAI, LlamaIndex, Kubernetes, FastAPI
Responsibilities
Turn ambiguous ideas into production systems across LLM features, agent workflows, retrieval, and memory; Develop async Python services supporting streaming and real-time interactions; Design short-term, long-term, and episodic memory for persistent conversational state; Build evaluation datasets, identify regressions, and improve quality, latency, and cost; Document decisions and take ownership of work from design through delivery