AI Specialist
Must have (6+ years, of which 3+ in production LLM systems)
Production deployment of open-source LLMs — Llama, Falcon, Jais, Mistral, or equivalent — on self-hosted GPU infrastructure
vLLM, Text Generation Inference, or equivalent high-throughput serving frameworks
RAG architecture in production: chunking strategies, embedding models, retrieval ranking, reranking, query rewriting
Vector databases — pgvector, ChromaDB, Weaviate, Qdrant, or Milvus (pgvector and ChromaDB are the target stack)
Embeddings for bilingual / multilingual content with RTL handling
Fine-tuning and parameter-efficient fine-tuning (LoRA, QLoRA, PEFT) on transformer models
LLM evaluation frameworks — RAGAS, TruLens, or equivalent — and ability to design custom evaluation harnesses
Python, PyTorch, Hugging Face Transformers, LangChain or LlamaIndex
NVIDIA GPU operations — CUDA awareness, multi-GPU inference, NVLink topology, model parallelism
Strongly preferred
MCP (Model Context Protocol) server development experience
Experience with UAE-origin models — Falcon, Jais, Jais Chat — or close equivalents
Arabic NLP — tokenisation, normalisation, dialectal handling (Gulf, Saudi, MSA), diacritic management
Prompt engineering at production scale , including chain-of-thought, few-shot, and instruction templates
LLM guardrails — Guardrails AI, NeMo Guardrails, or equivalent — for output filtering, PII redaction, jailbreak prevention
Kubernetes operation of GPU workloads , including NVIDIA k8s operators
Document AI — Puppeteer or equivalent for PDF rendering with bilingual RTL/LTR typography
Position details
Location: United Arab Emirates
Industry: Business Intelligence, Big Data & Analytics
Experience required: 8.0 years