Data Engineer (Databricks)

AcaisoftWarszawanofluffjobsопубліковано 02.09.2026
Обов'язково:PythonCloudDataCI/CDAIJunior

Hi there! If you’re looking for a high-impact position in an ambitious software house, we’ve got a match for you!

Currently, we are searching for a full-time Data Engineer to cooperate with our US client - BioPharma organization. The mission is to unlock the potential of AI/ML to improve the lives of patients. Our client delivers powerful cloud software and services specifically designed to meet the evolving needs of the BioPharma sector.

You will b e an integral part of the hands-on delivery of reliable, scalable data pipelines and datasets that power AI discoverability, analytics, reporting, and workflow automation across scientific, clinical, and enterprise data.

⚠️ Due to the client's location, the team is working till 6:00 PM CEST.

Daily tasks

  • Build, evolve, and operate data ingestion and processing capabilities for structured, semi-structured, and unstructured data, supporting the transition from early prototypes through early adoption and general release.
  • Implement and maintain rich metadata and data quality practices that enable cross- record querying, traceability, and AI-ready data access across experiments, files, inventory, and workflows.
  • Work with architects, AI engineers, workflow engineers, and domain experts to ensure data is usable, performant, and trustworthy for downstream GenAI and analytics use cases.
  • Support engineering quality through code reviews and mentoring of junior/mid-level engineers.

Requirements

Strong hands-on Python and complex SQL, with production Spark experience on Databricks (Structured Streaming, Auto Loader, Delta Lake, Unity Catalog) to support analytics, AI/ML, and enterprise application use cases. Practice working with structured and unstructured data across the full lifecycle: ingestion, transformation, enrichment, indexing, and retention, delivered through disciplined engineering practice including infrastructure-as-code, CI/CD, and automated testing of pipelines. Proven experience ingesting from heterogeneous operational sources including relational (Oracle, PostgreSQL) and document stores (MongoDB), using Full load + CDC or equivalent replication patterns, with practical handling of schema drift, deletes, and late or out-of-order data. Skills in data modelling across a layered architecture, with clear contracts between raw, curated and serving layers, supporting cross-entity queries and contextual linking, and performant data access patterns. Nice to have: Experience building unstructured document pipelines: parsing and extraction from PDF, Office and scanned formats, chunking, embedding generation, and maintaining searchable indexes at scale.

Must have: Python, SQL, Spark, AI, PostgreSQL, MongoDB, Data modelling