Data Engineer (Apache Spark)

Aptivamena TechCairowuzzufveröffentlicht 13.08.2026
Muss:PythonJavaDockerKubernetesCloudBackendDataAI

Build high-scale batch and near-real-time data pipelines deployed on infrastructure we run ourselves (on-prem), not managed cloud services. You will design and operate high-volume analytical data systems end to end, with Apache Spark as the core processing engine for both batch and streaming workloads.

7+ years of experience in data engineering and software development

Ability to write high-quality code in Java/Scala, Python or equivalent languages

Deep, hands-on production experience with Apache Spark - batch and Spark Structured Streaming (core requirement)

Demonstrated Spark performance tuning: partitioning, caching and persistence, broadcast joins, shuffle reduction, data-skew handling and Adaptive Query Execution

Experience operating Spark on self-managed clusters (YARN, Kubernetes or standalone) - executor sizing, resource allocation and multi-tenant workloads

Practical experience with Kafka (or equivalent messaging systems) as a Spark source and sink for high-volume workloads, including offset and checkpoint management

Practical experience with distributed query engines (e.g. Trino/Presto or similar)

Practical experience with ETL / data integration tools (e.g. Datastage, Informatica, Apache NiFi) and SQL-based transformation frameworks (e.g. dbt)

Strong SQL skills and understanding of data modeling and data warehousing for analytical workloads

Hands-on experience with real-time / low-latency analytical stores (columnar or OLAP engines, e.g. Apache Pinot/ClickHouse or similar)

Practical experience with big-data platforms and distributions (e.g. Cloudera, Hadoop ecosystem, Databricks or similar)

Practical experience containerizing and operating data workloads (Docker; Kubernetes a plus) and workflow orchestration tools (e.g. Airflow)

Familiarity with data lake table formats (e.g. Apache Iceberg, Delta Lake or similar), including schema evolution and compaction

Familiarity with data governance / cataloging tools (e.g. DataHub) and lakehouse management systems (e.g. Apache Amoro)

Familiarity using AI tools for development and debugging (Claude, Cursor, Codex)