AI Inference Engineer QVAC (100% remote Worldwide)
Core
Own the C++ inference backbone for QVAC's local AI stack, enabling fast, reliable, and private on-device AI experiences without cloud reliance.
Role type
Senior IC C++ inference engineer (edge AI)
Builds
Local AI inference engines and runtime systems for QVAC's peer-to-peer products
Domain
Fintech / Edge AI / On-device ML
Deliverable
production ML models
Required skills
C++, llama.cpp, ggml, GPU programming (CUDA/Vulkan/Metal/OpenCL), deep learning concepts, transformer/LLM/diffusion model architectures
Preferred skills
LLM training/fine-tuning, productionized models, new model architecture research, distributed systems, JavaScript
Technologies
llama.cpp, ggml, CUDA, Vulkan, Metal, OpenCL
Responsibilities
Deploy machine learning models to edge devices using llama.cpp and ggml; Collaborate with researchers to transition models from research to production; Integrate AI features into existing products
Seniority
Senior, hands-on IC