Senior Data Product Engineer Talend / Starburst - Freelance (M/F)
Must-have:PythonKubernetesDevOpsDataAgileCI/CD
Machine translation — original language: French.Show original
CONTEXT
The client is seeking support to ensure a mission within the framework of its transition towards a Data Mesh architecture. We are looking for a Data Product Engineer specialized in developing pipelines on the Talend Real Time BigData platform and data products on the Starburst platform.
MISSIONS
- Design of data flows and Data Products
- Development of optimized data integration pipelines via Talend Data Fabric / Talend Real Time BigData or Spark / Astronomer
- Documentation of data ingestion pipelines
- Integration of data into the Data Mesh ecosystem
- Collaboration with business domains to identify data sources and use cases
- Development of data products / data sets (views & materialized views) meeting use cases identified with business domain managers
- Development of optimized SQL queries for Starburst (push-down predicates, partitioning, file formats like Iceberg/Parquet)
- Documentation of data products (metadata, ownership, quality)
- Contribution to the continuous improvement of data products within the Starburst platform
- Implementation of alerts and dashboards for monitoring data products
TOOLS & ENVIRONMENT
- Apache Spark
- Talend (Talend RealTime BigData, Talend Studio)
- Starburst / Trino
- Airflow / Astronomer
- Storage formats: Iceberg, Parquet, S3
- Governance tools: Ranger, RBAC
- DevOps/DataOps practices: CI/CD, infrastructure-as-code (Terraform, Kubernetes)
- Languages: advanced SQL, Python (Pandas, PySpark), bash
WORKING CONDITIONS
- Mission location: Paris
- Start date: ASAP
- Long-term mission
- Remote work: 2 days per week
- Professional English mandatory
- Required seniority: 8 to 10 years minimum experience
- 8 to 10 years minimum experience
- Advanced experience in developing pipelines on Talend RealTime BigData (Talend Studio in batch or route mode)
- Advanced experience in SQL on Starburst/Trino (joins, CTEs, analytical functions)
- Data Modeling skills
- Experience with Airflow/Astronomer, Spark, Talend Real Time Big Data / Talend Data integration
- Knowledge of Iceberg, Parquet, S3 storage formats
- Familiarity with security and governance tools (Ranger, RBAC)
- CI/CD practices for pipelines, infrastructure-as-code (Terraform, Kubernetes)
- Mastery of languages: advanced SQL (mandatory), Python (Pandas, PySpark), bash
- Understanding of Data Mesh principles (domain-oriented ownership, self-serve platform, federated governance)
- Ability to work in product mode (roadmap, prioritization, collaboration with business units)
- Sensitivity to data quality (tests, monitoring, remediation) and documentation (data contracts, metadata)
- Collaborative spirit: working with business teams, data scientists, and DevOps
- Pedagogy: ability to explain technical concepts to non-experts
- Autonomy: taking initiative within an agile framework