Senior Data Quality Engineer
Job purpose
At Emirates Group our Analytics Centre of Excellence (ACoE) is a centralized unit that provides data and analytics support to Emirates Group businesses. This allows our businesses to make better decisions by using data and analytics to understand our customers, operations, and markets. The ACoE is an essential part of Emirates Group's digital transformation strategy. The unit is helping Emirates Group to become a more data-driven organization and to make better decisions by using data and analytics.
As a senior data quality and automation engineer you will be a fully participating member of a cross-functional team working autonomously on technology development and problem resolution in the enterprise data and analytics space. The role involves hands-on contributing to quality practices along with designing and implementing data quality and automation platforms. You will also provide support to technical analytics solutions and products that support Emirates Airlines and the Emirates Group businesses.
In this role you will:
Work closely with product owners, analysts, software engineers and architects to understand the technical landscape and context of deliveries to refine complex functional and non-functional requirements and translate it into fit for purpose acceptance tests.
Work on problems of diverse scope where analysis of data requires evaluation of identifiable factors and select methods to validate and automate the solution.
Build test strategies and test plans validating the core business problem and translating them into tests around code functionality, data quality, performance, and security.
Build and support generating or mocking of test data for exploratory analysis and running tests.
Design and build automated tests and jobs of moderate to high scope and complexity across all stages of the data pipeline while demonstrating good coding and design practices and adhering to published coding standards and guidelines.
Build and enhance data quality rules and platforms for observability across the data pipeline.
Debug complex issues, resolve blockers and follow design documents with minimal or no supervision.
Conduct data analysis activities such as source system analysis, data modelling, data dictionary collection, data profiling and source-to-target mapping ensuring delivery on business needs.
Support in updating data inventories and registries as required to keep metadata and data lineage up-to-date, following agreed data governance standards, guidelines and principles.
Qualification
To be considered for the role, you must meet the below requirements:
Qualifications and experience:
Degree in a relevant field such as computer science, computational mathematics, computer engineering or software engineering.
Specialization or electives in a data and analytics field (e.g. data warehousing, data science, business intelligence) is a nice-to-have.
Experience 2+ years data engineering experience with focus on quality and automation.
Minimum 2+ years of testing, automation and support experience in analytics applications such as data lake and data warehouse (preferably using the big data stack and Microsoft Azure cloud infrastructure).
Seasoned in identifying quality issues across complex data pipelines running on big data technologies and defining rules for validating health of the data.
Experienced in a wide variety of testing methods and tools covering functional, performance, security tests across individual jobs, pipelines and end to end across enterprise.
Well versed in building automated checks running in a CI pipeline to validate the ETL and ELT jobs are performing as expected.
Experience with batch and real-time data ingestion and integration tools and technologies handling massive quantities of data (structured and unstructured).
Exposed to data architecture concepts such as data modelling, big data storage, and dimensional modelling.
Exposed to working with jobs in data pipelines and define metrics and measures to ensure correctness of the data.
Programming (Python or Scala) and SQL querying skills are required.
Exposure to Spark and airline industry experience is nice-to-have.
Knowledge and skills:
Ability to drive quality of data assets independently.
Able to deliver solutions (and associated value) interactively.
Strong ability to conduct data analysis (e.g. source system identification, data dictionary and metadata collection, data profiling, source-to-target mapping) is preferred.
Operates with a "You code it, you own it" mindset (i.e. supports the products they build).
Team player; able to collaborate with others to remove blockers, solve complex design problems and debug and resolve issues.
Is accountable and displays positive attitude.
Self-starter and has passion for exploring and learning new technologies, especially those in the enterprise data and analytics space.
Technology domain key technologies and tools:
Big data and distributed processing: Spark, Hadoop (HDFS, Hive, HBase, Oozie), Airflow, Apache Nifi, Azure (ADLS, DataBricks, Azure Data Factory), Elasticsearch, AVRO and PARQUET file formats.
Data analysis, modelling and reporting: Snowflake, SQL, Data Vault 2.0, MicroStrategy, Power BI.
Cloud technologies: Microsoft Azure and Cloudera technology stacks.
Integration and messaging: Streaming (e.g. Spark Streaming), SnapLogic, TIBCO, Kafka.
CI/CD: GIT, Bitbucket, Jenkins, Azure DevOps, Kubernetes, Docker, SonarQube, Gatling.
Languages: Scala, Python.
Salary and benefits
Join us in Dubai and enjoy an attractive tax-free salary and travel benefits that are exclusive to our industry, including discounts on flights and hotel stays around the world.