Principal ML Engineer

RYZ LabsArgentinaJob.bopublished 10/09/2026
Must-have:TypeScriptPythonAWSAzureGoogle CloudDockerCloudBackendDevOpsCI/CDMicroservicesAISecurityPrincipalRemoteHybrid

Location: Remote, only for candidates in LATAM

Job Overview

Ryz Labs is looking for a Principal Machine Learning Engineer to join the core engineering team of one of our key international clients. In this role, you will play a pivotal part in shaping our client’s next-generation AI platforms, sitting at the intersection of production systems, applied AI, and engineering excellence.

You will bridge distributed systems architecture with hands-on MLOps and LLM engineering. We are seeking an exceptional technical leader who can design scalable multi-tenant architectures, build robust AI agent governance frameworks, and serve as a trusted technical authority collaborating directly with product managers and key client stakeholders.

Key Responsibilities

  • Production ML & Optimization: Deploy and manage AI models at scale, monitoring performance, hallucination rates, drift, latency, and infrastructure costs.
  • Architecture & Delivery: Design distributed, event-driven microservices using Python, Go, or TypeScript while building IaC and CI/CD pipelines to ship your own services.
  • AI Security & Governance: Implement agent permission structures, human-in-the-loop workflows, data isolation boundaries, and prompt injection defenses.
  • Product Collaboration: Partner with Product Managers from inception to translate business requirements into scalable architectures and present trade-offs to executives.

What You Bring

  • 12–15+ years in software engineering, with a clear evolution from Backend/Distributed Systems Architecture into applied Production ML Engineering.
  • Production ML Expertise: Deep experience with MLOps, model evaluation rubrics, advanced RAG, vector search (embeddings, HNSW, hybrid search), and fine-tuning. (We are looking for engineers building real systems, not just consuming LLM APIs).
  • Software Architecture: Strong mastery of distributed systems, microservices, and asynchronous event-driven patterns in Python, Go, or TypeScript.
  • Practical DevOps & Cloud: Hands-on command of Docker, cloud infrastructure (AWS/GCP/Azure), and automated CI/CD pipelines.
  • AI Governance & Security: Practical knowledge of LLM safety, threat modeling, data boundary enforcement, and agent security.
  • Fluent English & Communication: Ability to articulate complex technical trade-offs (e.g., RAG vs. Fine-tuning, latency vs. accuracy) clearly to client executives and non-technical stakeholders.

Nice-to-Have

  • Prior experience with multi-agent orchestration frameworks (e.g., LlamaIndex, Semantic Kernel, AutoGen, CrewAI, MCP).
  • Experience in fast-paced consulting, advisory, or high-growth tech platforms.