Senior Machine Learning Systems Engineer
Accountabilities Design and implement end-to-end model lifecycle and MLOps patterns covering data preparation, model management, experiment tracking, deployment, and related workflows.
Develop and support graph machine learning platforms and codebases that abstract common patterns and enable scalable model development and iteration.
Partner with machine learning engineers to improve training performance, reduce model training times, increase efficiency, and optimize GPU costs across distributed environments.
Optimize large-scale batch data processing using cloud data warehouses and distributed processing technologies such as Apache Beam, Apache Spark, and Ray Data.
Architect pipelines capable of creating and maintaining massive graph datasets containing billions of nodes and tens of billions of edges.
Administer and integrate tooling for experiment tracking, model serving, model registries, and other components of the ML platform.
Build infrastructure and platform capabilities that prioritize scalability, reliability, performance, maintainability, and ease of use.
Collaborate closely with platform users to understand their development lifecycle and remove technical friction from machine learning workflows.
Contribute to architectural and technical decisions across cloud infrastructure, distributed training, data processing, and machine learning systems.
Help improve platform efficiency and cost management while maintaining reliable infrastructure for large-scale ML workloads.
Requirements
5+ years of professional experience working with machine learning infrastructure, including model training and deployment.
Hands-on experience optimizing machine learning workloads, including memory utilization, GPU profiling, training performance, and resource efficiency.
Deep experience with cloud technologies supporting ML platforms, particularly services such as GCP BigQuery and Google Cloud Storage, as well as infrastructure-as-code tools such as Terraform.
Practical experience administering and integrating MLOps technologies for experiment tracking, model serving, and model registries, such as MLflow or Weights & Biases.
Strong programming skills in languages and frameworks commonly used for machine learning, including Python, PyTorch, and/or TensorFlow.
Deep experience with distributed training and compute frameworks, including Ray and Kubernetes.
Strong understanding of scalable data processing and distributed systems, with the ability to work effectively across complex ML infrastructure environments.
A strong focus on scalability, reliability, performance, and developer experience, with an ability to understand and advocate for the needs of platform users.
Strong organizational, communication, analytical, and problem-solving skills, with the ability to collaborate effectively across technical teams.
Experience with graph databases such as Neo4j, JanusGraph, or TigerGraph is a plus.
Experience with graph neural networks and graph ML frameworks such as PyTorch Geometric or Deep Graph Library is a plus.
Benefits
Base salary range of $216,700–$303,400 USD , with final compensation determined by factors including skills, experience, credentials, and other job-related considerations.
Eligibility for equity in the form of restricted stock units.
Potential eligibility for commission depending on the position offered.
Medical, dental, and vision insurance options for U.S.-based employees.
401(k) program with employer matching.
Generous paid time off and vacation benefits.
Parental leave to support employees and their families.
Remote work from within the United States.
Opportunities to work on large-scale machine learning infrastructure, distributed systems, graph ML, and cloud technologies.
A technically collaborative environment focused on innovation, scalability, reliability, and continuous improvement.
How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1