Big Data Engineering Lead
Must-have:PythonAWSCloudBackendDataAgileCI/CDAILead
Job Description
We are seeking an experienced Big Data Engineering Lead to design, develop and lead scalable enterprise data platforms, real-time data pipelines and analytics solutions. The successful candidate will work with large-volume datasets, cloud platforms, distributed processing technologies and modern AI/GenAI solutions.
Key Responsibilities
- Lead the design and development of scalable Big Data Engineering and Data Analytics platforms .
- Design and implement batch and real-time data pipelines using Python, PySpark, Spark Structured Streaming and Kafka .
- Develop data ingestion, transformation and processing workflows using AWS, S3, Airflow and orchestration tools .
- Design data models, data warehouses, data lakes/lakehouse architectures and high-performance analytics solutions.
- Develop and optimize complex SQL queries, ETL processes and data processing workflows .
- Build REST APIs and backend services using FastAPI/Flask for data and analytics applications.
- Implement data replication, archival, reconciliation, quality monitoring and governance processes.
- Deploy scalable data applications on AWS , including containerized environments and CI/CD pipelines.
- Lead development of GenAI/LLM, RAG and Agentic AI solutions for enterprise analytics and natural-language data querying.
- Provide technical leadership, mentoring and guidance to data engineering teams.
- Collaborate with business and technology stakeholders to translate requirements into scalable data solutions.
Requirements
- Degree in Computer Science, Information Technology, Engineering, Data Science or a related discipline.
- 8+ years of experience in Data Engineering / Big Data / Data Analytics.
- Strong hands-on experience with Python, SQL, PySpark, Apache Spark and Kafka .
- Experience in AWS cloud data services , including S3 and related data engineering technologies.
- Strong experience with Airflow or equivalent workflow orchestration tools .
- Experience with Snowflake, Hadoop/HDFS, Hive or other enterprise data platforms .
- Strong understanding of ETL, data warehousing, data modelling and distributed data processing .
- Experience developing REST APIs using FastAPI or Flask .
- Exposure to GenAI, LLM, RAG, LangChain/LangGraph or Agentic AI is an advantage.
- Strong analytical, problem-solving and technical leadership skills.
- Experience working in Agile software development environments.