Data Integration Engineer
Must-have:PythonJavaGitAWSDockerKubernetesCloudDataCI/CDAISenior
Job Summary
We seek a Senior Data Engineer to design and optimize scalable data pipelines and processing solutions using Big Data, cloud, and AI technologies. Collaborate with architects, data scientists, and engineers to deliver high-performance, production-ready data systems.
Responsibilities
- Design, develop, and maintain scalable batch and real-time data pipelines using Apache Spark/PySpark, SQL, and Python to support large-scale data processing
- Develop and optimize ETL/ELT pipelines to ensure efficient data ingestion and transformation
- Build data ingestion frameworks integrating databases, APIs, files, and streaming platforms for reliable data flow
- Utilize AWS services including S3, Glue, EMR, Redshift, Kinesis, Lambda, and DynamoDB to deploy and manage cloud-based data solutions
- Develop and optimize data processing workflows using Hadoop, Hive, Spark, Kafka, Cloudera, and Databricks for high throughput and low latency
- Perform SQL and Spark performance tuning to enhance processing speed and resource utilization
- Design and implement data models, data warehouses, and data lake architectures to support analytics and reporting needs
- Develop CI/CD pipelines and manage containerized deployments using Docker, Kubernetes, and OpenShift to streamline development and operations
- Integrate Generative AI and NLP capabilities into enterprise data applications to enhance data insights and automation
- Develop solutions leveraging LLM frameworks, RAG, vector databases, and AI APIs to support advanced AI-driven data applications
- Collaborate with solution architects, data scientists, software engineers, and business stakeholders to deliver scalable, production-ready data solutions
- Participate in system design, development, testing, deployment, and provide production support to ensure system reliability and performance
Required competencies and certifications
- Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or related field
- Minimum 8 years of experience in Data Engineering / Big Data with expertise in large-scale data processing
- Proficient programming skills in Python, Java, and/or Scala
- Hands-on experience with Apache Spark/PySpark, SQL, Hadoop, and Hive
- Experience working with AWS cloud platform and its data services
- Knowledge of data warehouses, relational databases, and data lake technologies
- Experience with Kafka, Airflow, Jenkins, Git, and CI/CD practices
Preferred competencies and qualifications
- Experience with Docker, Kubernetes, or OpenShift container orchestration platforms
- Familiarity with Databricks, Snowflake, and Cloudera platforms
- Exposure to Generative AI, LLMs, RAG, LangChain/LangGraph, and vector databases
- Strong analytical, problem-solving, and communication skills