AI Platform & LLMOps Expert (Principal / Lead)
About the Project / Role Context: For our client – an international FinTech & Telecom group – we are building a world-class AI capability (AI Squad) from the ground up. We create platforms, agents, and intelligent systems that will power the next generation of financial services across multiple markets. We are seeking an AI Platform & LLMOps Expert to build, operate, and maintain the platform on which every AI agent, model, and application is deployed, evaluated, monitored, and run. This is a hands-on platform engineering and technical leadership role – you will own the deployment pipelines, evaluation harness, observability stack, model routing, cost controls, and user-facing APIs/interfaces.
Daily tasks
- AI Platform & Release Mechanics: Build and maintain multi-environment AI platform infrastructure, CI/CD pipelines, container orchestration, and release mechanics.
- Model Operations & LLMOps: Manage production model lifecycle, including provisioning, dynamic routing, versioning, fallback strategies, latency optimization, and cost control.
- Evaluation & Quality Gates: Build and own evaluation harnesses, regression test suites, and pre-release quality gates to ensure model/agent safety and performance.
- Observability & Incident Leadership: Design the full observability stack (monitoring, tracing, alerting) and hold production accountability (on-call, incident response, root-cause analysis).
- Application Interfaces & APIs: Deliver full-stack applications, developer interfaces, and robust APIs through which services and end-users consume AI capabilities.
Requirements
Requirements (Must-have): Experience: Minimum 8–10 years in software, platform, cloud, or Site Reliability Engineering (SRE), including 2–3 years leading or guiding engineering teams . Production Platform Engineering: Proven hands-on experience building and running enterprise-grade cloud platforms, IaC, containers, and automated pipelines. LLMOps / MLOps Practice: Practical experience operating LLM or ML workloads in production, with a focus on model evaluation, monitoring, tracing, and token/infrastructure cost management. Cloud & Infrastructure: Strong proficiency with cloud platforms (AWS/Azure/GCP), Kubernetes/containers, Infrastructure as Code (Terraform/Ansible), and secrets management. Production Accountability: Demonstrated experience with on-call shifts, incident response leadership, and root-cause analysis in high-availability environments. Regulated Environments: Exposure to regulated sectors with formal change management, audit, and security requirements. Nice-to-have: Recognized cloud, platform, or SRE certifications (e.g., AWS/Azure DevOps/Solutions Architect, CKA, TOGAF). Postgraduate qualification (Master's degree) in a technical field. Experience in FinTech, banking , or other high-volume transactional platforms.
Must have: LLMOps, MLOps,, Cloud, Kubernetes, Terraform, Python, CI/CD, Observability
Nice to have: DevOps/Solutions Architect,, TOGAF