Data Engineer #4885
Our mission is to detect cancer early, when it can be cured. We are working to change the trajectory of cancer mortality and bring stakeholders together to adopt innovative, safe, and effective technologies that can transform cancer care.
We are a healthcare company, pioneering new technologies to advance early cancer detection. We have built a multi-disciplinary organization of scientists, engineers, and physicians and we are using the power of next-generation sequencing (NGS), population-scale clinical studies, and state-of-the-art computer science and data science to overcome one of medicine’s greatest challenges.
GRAIL is headquartered in the bay area of California, with locations in Washington, D.C., North Carolina, and the United Kingdom. It is supported by leading global investors and pharmaceutical, technology, and healthcare companies.
For more information, please visit grail.com
As a Senior Data Engineer on the Operational Technology team, you will own the data platform that connects GRAIL's lab instruments, automation systems, and operational platforms to a trusted, well modeled data foundation. You will architect and lead the development of complex, business-critical ingestion and transformation pipelines end to end, set the standards and patterns the broader team builds on, and act as a technical point of contact across systems engineering, lab operations, data science, and automation engineering. You will work independently on problems of diverse scope, devising solutions where precedent is limited, and you will mentor less experienced engineers while raising the bar on reliability, data quality, and engineering practice. This is a hands-on senior role for a fully qualified data engineer who is ready to take ownership of critical infrastructure in a fast paced, regulated environment. Expect to work alongside a talented and highly motivated team that moves quickly. This role is based on-site in RTP, North Carolina, Monday through Friday. The position participates in an on-call rotation and may occasionally require weekend or holiday support for production incidents, maintenance, or critical deployments.
Responsibilities: Architect, build, and maintain complex data pipelines that ingest and integrate information from laboratory instruments, automation systems, sequencers, operational platforms, APIs, autonomous robotics platforms, databases and file based data sources.
Own critical pipelines and data models end to end, from design through production operation, working independently on problems of diverse scope and adapting existing approaches where limited precedent exists.
Define and evolve the data architecture, modeling standards, and engineering patterns that the broader team builds on, and drive adoption across the group.
Support downstream analytics, reporting, and AI systems by delivering clean, trustworthy datasets and timely data extracts for troubleshooting, root-cause investigations and platform improvements.
Develop and optimize advanced SQL and transformation logic to cleanse, standardize, and model raw instrument and production data into reliable, well structured datasets.
Build and support datasets and data models used by operational dashboards, analytics, process monitoring, troubleshooting, and governed AI enabled workflows.
Design and implement orchestration, testing, monitoring and alerting so that data failures, freshness issues, schema changes, and incomplete processing are identified and resolved early.
Establish and enforce data validation and quality standards to ensure datasets are accurate, complete, and reliable across the platform.
Partner with and advise systems engineers, lab operations, data scientists, and automation engineers on difficult technical matters, adapting your communication for both technical and non-technical stakeholders.
Mentor and provide technical guidance to junior engineers, and contribute to the team's overall engineering practices and standards.
Document pipelines, data models, and datasets to support reproducibility and compliance with ISO, CLIA, CAP, NYS, GMP, and FDA requirements.
Continuously improve your technical skills and the team's engineering practices.
Required Qualifications: Degree in Computer Science, Mathematics, Software Engineering, Data Science, Life Sciences, Physics or similar field.
Typically 3+ of relevant professional experience in data engineering, analytics engineering, or software development with a BS/BA degree; or 2+ with a Master's degree; or equivalent practical experience.
Advanced proficiency in SQL, including performance optimization and complex transformation logic.
Strong proficiency with one or more programming languages, such as Python, Rust, C++, or similar.
Solid understanding of ETL/ELT pipeline design, relational databases, data modeling, and structured or semi-structured data.
Demonstrated ability to own data pipelines and infrastructure end to end and to work independently on problems of diverse scope.
Strong attention to detail and a commitment to data quality, reliability and accuracy.
Ability to collaborate effectively with, and advise, both technical and non-technical individuals, and comfort working in a rapidly changing environment with dynamic objectives and fast iteration.
Ability to investigate complex technical problems methodically, continuously learn, and communicate clearly to a range of audiences.
A highly analytical mindset and eagerness to solve difficult technical problems.
Preferred Qualifications: Hands-on experience with data pipeline orchestration and transformation tools such as Airflow, dbt, or comparable technologies.
Experience with cloud data platforms, object storage and warehouses such as AWS S3, Redshift, Glue, Snowflake or comparable technologies.
Experience integrating AI/agentic tooling into the data engineering SDLC.
Experience with semantic data modeling, data lineage, and automated data quality testing.
Experience mentoring engineers or leading technical projects and setting engineering standards.
Familiarity with statistical methods or process analytics.
Exposure to manufacturing, clinical laboratory operations, diagnostics, or biotechnology.
Proficiency with version control systems such as Git and collaborative development practices.
Understanding of APIs, file transfers, networking and system integrations.
The expected, full-time, annual base pay scale for this position is 86K - $106K. Actual base pay will consider skills, experience, and location.
This role may be eligible for other forms of compensation, including an annual bonus and/or incentives, subject to the terms of the applicable plans and Company discretion. This range reflects a good-faith estimate of the range that the Company reasonably expects to pay for the position upon hire; the actual compensation offered may vary depending on factors such as the candidate’s qualifications. Employees in this role are also eligible for GRAIL’s comprehensive and competitive benefits package, offered in accordance with our applicable plans and policies. This package currently includes flexible time-off or vacation; a 401(k) retirement plan with employer match; medical, dental, and vision coverage; and carefully selected mindfulness programs.
GRAIL is an equal employment opportunity employer, and we are committed to building a workplace where every individual can thrive, contribute, and grow. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, sex, gender, gender identity, sexual orientation, age, disability, status as a protected veteran, , or any other class or characteristic protected by applicable federal, state, and local laws. Additionally, GRAIL will consider for employment qualified applicants with arrest and conviction records in a manner consistent with applicable law and provide reasonable accommodations to qualified individuals with disabilities. Please contact us at rc@grailbio.com if you require an accommodation to apply for an open position.
GRAIL maintains a drug-free workplace. We welcome job-seekers from all backgrounds to join us!