Principal ML System Engineer
Accountabilities: Define and drive the technical vision and multi-quarter roadmap for the organization’s machine learning platform, partnering with Product and Engineering leadership to align platform investments with business and product objectives.
Establish reference architectures, engineering standards, and scalable patterns for data and ML pipelines covering model training, evaluation, deployment, and serving.
Set company-wide MLOps practices, including model CI/CD, model registries, feature stores, experiment tracking, and other foundational capabilities, while leading build-versus-buy decisions for platform components.
Design and establish reliability, observability, and performance practices for production ML systems, including monitoring, alerting, automated remediation, and operational standards.
Define the ML platform’s security architecture, including authentication, role-based access control, audit logging, compliance monitoring, and secure adoption patterns across engineering teams.
Develop secure, scalable, and cost-efficient infrastructure and integration patterns connecting ML capabilities with existing systems, APIs, and enterprise data sources.
Provide technical leadership across multiple engineering teams, influencing organization-wide architecture decisions and guiding senior engineers on complex ML infrastructure challenges.
Lead and sustain the development of critical cross-team ML systems, ensuring they remain scalable, maintainable, reliable, and aligned with evolving business needs.
Support strategic decisions involving cloud platforms, ML frameworks, data processing environments, model-serving technologies, and emerging AI infrastructure.
Drive optimization of large-scale model training and inference workloads, balancing performance, reliability, scalability, and infrastructure costs.
Promote strong software engineering, security, and operational practices across the ML platform and contribute to raising the overall engineering standard.
Requirements:
Expert-level proficiency in Python and Java , combined with strong software engineering fundamentals and experience building production-grade systems.
Deep experience designing and implementing ML platforms and MLOps workflows at scale , with familiarity with technologies such as MLflow, Kubeflow, Ray, model-serving frameworks, or equivalent platforms.
Extensive experience working with major cloud platforms such as AWS, Azure, and/or GCP , as well as containerization and orchestration technologies including Docker and Kubernetes.
Demonstrated ability to establish technical direction and successfully drive complex, organization-wide engineering initiatives spanning multiple teams.
Strong understanding of machine learning lifecycle architecture, including training, evaluation, deployment, serving, monitoring, experimentation, and operational management.
Experience designing secure systems at scale, including role-based access control, multi-factor authentication, network security practices, auditability, and compliance monitoring.
Strong knowledge of distributed systems, production reliability, observability, scalability, and performance optimization for ML workloads.
Familiarity with Azure Machine Learning, Databricks, serverless environments, and modern ML frameworks is preferred.
Experience leading the development and long-term operation of critical cross-functional or cross-team technology platforms.
Experience optimizing large-model training and inference, including LLM serving , for performance and cost efficiency is an advantage.
Bachelor’s degree or higher in Computer Science, Machine Learning, or a related discipline is preferred.
Excellent communication, influence, and mentoring skills, with the ability to explain complex technical concepts and shape decisions across engineering and leadership audiences.
Benefits:
Competitive base salary of $176,000–$195,000 per year , with compensation aligned to experience, skills, and market context.
Additional bonus and benefits as part of the overall total rewards package.
Fully remote work option in Canada , with the opportunity to work from Mississauga.
Opportunity to shape organization-wide machine learning and generative AI infrastructure strategy .
High-impact technical leadership role with broad influence across engineering teams.
Exposure to advanced ML, MLOps, cloud infrastructure, LLM serving, security, and large-scale AI systems.
Opportunities to mentor senior engineers and contribute to engineering standards and best practices.
Collaborative environment involving Product, Engineering, and cross-functional technical partners.
Opportunity to lead the development of scalable, secure, reliable, and cost-efficient AI capabilities.
How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1