Machine Learning Engineer (Ops)
You will build trusted data and trusted AI - ensuring our clients data is accurate, compliant, and governed, and our ML models are reproducible, monitored, and responsibly deployed to production. This role is 50% Data Governance, 50% MLOps / ML Platform Governance. Key Responsibilities A. Data Governance (50%) Framework & Stewardship Design and run enterprise Data Governance framework, policies, and RACI for data owners/stewards Establish Data Governance Council and operating model across Product, Engineering, Analytics, and Business Define KPIs: catalog coverage, data quality score, policy adherence
Data Quality, Catalog & Lineage Implement business glossary, data catalog (Collibra / Alation / Purview / DataHub), and end-to-end lineage Define and monitor data quality rules, SLAs, anomaly detection for critical domains (Customer, Product, Transaction) Manage data classification, PII/PHI tagging, retention, and access control policies
Compliance & Security Ensure compliance with PDPA, GDPR, CCPA and internal security standards Partner with DPO / Legal / GRC for consent, purpose limitation, anonymization, and audit readiness Own access governance - RBAC/ABAC for data warehouse, lakehouse, and feature store
B. MLOps & AI Governance (50%) ML Lifecycle & Platform Own MLOps best practices: from feature engineering -> training -> validation -> deployment -> monitoring Build and manage ML platform components: Feature Store (Feast / Tecton / SageMaker Feature Store), Model Registry (MLflow / SageMaker Model Registry), Experiment Tracking Standardize CI/CD/CT for ML with Git, Docker, Airflow / Kubeflow / SageMaker Pipelines
Model Governance & Responsible AI Implement Model Governance: model inventory, model cards, lineage (data -> features -> model -> endpoint), approval workflows Enforce responsible AI checks: bias/fairness, explainability, drift, and reproducibility Align with AI Governance frameworks: NIST AI RMF, Singapore Model AI Governance Framework, AI Verify, ISO 42001
Monitoring & Operations Implement monitoring for data drift, concept drift, feature skew, and model performance degradation Set up alerting, automated retraining triggers, and rollback strategies Optimize model serving costs, latency, and scalability on AWS / Azure / GCP
Tech Stack You Will Work With Governance: Collibra, Alation, Purview, Informatica, DataHub, AWS Glue, Apache Atlas Data: Snowflake / BigQuery / Redshift, S3 / GCS, dbt, Airflow, Spark, Kafka MLOps: MLflow, Kubeflow, SageMaker, Vertex AI, Feast, Evidently, Great Expectations, Docker, Kubernetes, GitHub Actions Languages: Python (must), SQL (must), PySpark Requirements 6-10 years total in Data Engineering / Data Governance / MLOps At least 2+ years owning data governance and at least 2+ years deploying ML models to production Strong hands-on with DAMA-DMBOK and MLOps principles Proven experience setting up Model Registry, Feature Store, and monitoring for production ML systems Deep understanding of PDPA/GDPR, data security, and AI risk Excellent stakeholder management - you can talk to both Data Scientists and Risk/Legal Nice-to-Have CDMP, AWS Certified ML Specialty, or similar Experience with LLM / GenAI governance - prompt logging, RAG governance, hallucination monitoring Experience with Great Expectations, Monte Carlo, Evidently AI Industry experience in Media, FinTech, or other regulated industry What Success Looks Like in 12 Months Top 5 data domains governed with SLAs and quality monitoring >95% 100% of production models registered with model cards, lineage, and approval workflow Automated drift detection live for all critical models with Data catalog adoption >80% and zero compliance audit findings