Data Engineer (Data Lake, Data Warehouse & MDM)
Must-have:PythonJavaAWSAzureGoogle CloudCloudDataScrumCI/CDTDDAIHybrid
About the Role
We are looking for an experienced Data Engineer to support enterprise data initiatives involving master data management, data governance, data lakes, data warehouses, and large-scale data processing platforms. The role focuses on building and enhancing data pipelines, managing complex datasets, and supporting data-driven initiatives across cloud, on-premises, and hybrid environments.
Key Responsibilities
- Create and manage a single master record for each business entity, ensuring data consistency, accuracy, and reliability.
- Implement data governance processes, including data quality management, data profiling, data remediation, and automated data lineage.
- Create and maintain multiple robust and high-performance data processing pipelines within cloud, private data centre, and hybrid data ecosystems.
- Assemble large, complex data sets from a wide variety of data sources.
- Collaborate with Data Scientists, Machine Learning Engineers, Business Analysts, and business users to derive actionable insights on data quality.
- Design and implement internal processes to automate manual workflows, optimize data delivery, and re-design infrastructure for greater scalability.
- Support and work with cross-functional teams in a dynamic environment.
Requirements
Experience
- At least 5 years of experience in a Data Engineer role.
- Experience building and operating large-scale data lakes and data warehouses.
- Experience working on projects for Master Data Management.
- Proven ability to support and work with cross-functional teams in a dynamic environment.
Data Platforms & Architecture
- Familiarity with data lake, data warehouse, and data lakehouse architectures.
- Familiarity with MDM processes such as golden record creation, survivorship, reconciliation, enrichment, and quality.
- Experience with Delta Lake and Databricks.
- Experience working with Hortonworks Data Platform or Cloudera Data Platform.
Data Governance & MDM
- Experience in data governance, including data quality management, data profiling, data remediation, and automated data lineage.
- Experience with Master Data Management (MDM) tools and platforms such as Informatica MDM, Talend Data Catalog, Semarchy xDM, IBM PIM & IKC, or Profisee.
- Experience with Metadata Management tools.
- Exposure to Data Governance processes and tools.
Big Data & Streaming Technologies
- Experience with Hadoop ecosystem and big data tools, including Spark and Kafka.
- Experience with stream-processing systems including Spark Streaming.
Database & Query Optimization
- Advanced working experience with relational SQL and NoSQL databases, including Hive, HBase, and Postgres.
- Deep understanding of SQL and the ability to optimize data queries.
- Successful experience manipulating, processing, and extracting value from large, disconnected datasets.
Programming & Development
- Experience with object-oriented and scripting languages such as Python, Java, and Scala.
- Experience applying modern development principles, including Scrum, Test-Driven Development (TDD), Continuous Integration (CI), Code Reviews
Cloud & ETL Technologies
- Experience working with cloud services using one or more cloud providers such as Azure, GCP, or AWS.
- Experience with ETL tools such as Talend Big Data, Azure Data Factory, or similar solutions.
Soft Skills
- Proven ability to support and work with cross-functional teams in a dynamic environment.