Senior MLOps / LLMOps Engineer
about job Monitor and maintain live Generative AI agents and Machine Learning models to ensure system accuracy, safety, and continuous uptime.
Design and implement operational observability dashboards to track performance, system health, and key business outcome metrics.
Build and execute continuous evaluation suites to proactively detect data drift, performance degradation, and quality issues in production.
Handle incident management and first-level technical triage, performing swift context-specific fixes or escalating with clear evidence logs.
Establish metrics and key telemetry signals alongside cross-functional engineering, architecture, and product management teams.
Participate in operational handover gates to verify system readiness and evaluation metrics prior to full production deployment.
skills and requirements Bachelor's degree in Computer Science, Engineering, Data Science, or a related field.
6 to 10 years of professional experience across Software Engineering, Site Reliability Engineering (SRE), Data Engineering, or ML Engineering.
Hands-on production experience managing and operating LLM applications, agentic frameworks, and AI system architectures.
Proficiency with LLM observability, tracing, evaluation tools/harnesses, and customized dashboarding solutions using Python.
Solid track record in incident triage, runbook execution, log/trace analysis, and real-time monitoring practices.
Strong technical communication skills to translate incident logs clearly and influence product build teams on overall quality standards. To apply online please use the 'apply' function, alternatively you may contact Sophia Lim at 9248 7488. (EA: 94C3609/ R26163753 )