Senior Data Engineer
Must-have:PythonJavaAWSAzureGoogle CloudDataAI
Responsibilities:
- Implement and operationalize enterprise Lakehouse platforms, data products, and data marketplace capabilities.
- Develop scalable batch, streaming, CDC, and API-based data ingestion pipelines.
- Develop scalable multimodal data ingestion pipelines including content extraction from various file formats, regex for specific field extraction, content extraction from embedded images, frame extraction from video files, transcript extraction from audio files, etc.
- Build, test, and maintain foundation and business data products with agreed data contracts, SLAs, and data quality controls.
- Implement open table formats such as Iceberg, Hudi, and Delta Lake.
- Support RAG, vector search, GenAI and agentic data pipelines.
- Perform performance tuning, optimization, production support, and root cause analysis.
- Create technical documentation, deployment guides, and operational runbooks.
- Ensure compliance with engineering standards, DevSecOps controls, and software delivery practices.
Requirement:
- 8-12 years of experience in Data Engineering, Big Data, Data Lake, or Lakehouse implementations.
- Hands-on experience with Databricks, Snowflake, Cloudera, Azure, AWS, GCP, Huawei, or Alibaba data platforms.
- Hands-on experience in developing Data products and Market place.
- Strong expertise in Spark, PySpark, SQL, Python and Scala.
- Strong programing skills (Java, Scala, Python, SQL).
- Experience with Iceberg, Hudi, Delta Lake and object storage platforms.
- Experience implementing data ingestion, transformation, reconciliation and data quality frameworks.
- Experience with Trino, Dremio, Hive, Impala, Kafka, Flink, Spark Streaming and Airflow.