Cloud Data Pipelines Engineer
Daily tasks
- R/Python Workload Execution: Containerize, deploy, and orchestrate custom R and Python scripts to run reliably in cloud-native serverless or containerized environments.
- Cloud Integration: Leverage managed cloud services on GCP (Cloud Functions, Cloud Run, BigQuery, Dataflow) or Azure (Azure Data Factory, Azure Functions, Databricks, Synapse Analytics).
- Workflow Orchestration: Implement and maintain workflow automation tools (such as Apache Airflow or Prefect) to schedule and monitor script execution.
- Data Quality & Reliability: Implement automated logging, error handling, and testing frameworks to validate data integrity and script outputs.
- Performance Optimization: Tune data processing jobs and cloud resource allocation for cost-efficiency, execution speed, and scalability.
Requirements
Programming: Advanced proficiency in Python and/or R , with strong experience writing clean, modular, and well-tested code. Cloud Platforms: Hands-on production experience with either Google Cloud Platform (GCP) or Microsoft Azure . Containerization & CI/CD: Practical experience with Docker , Kubernetes , and CI/CD pipelines (GitHub Actions, GitLab CI). Data Stores: Proficiency with SQL and relational/non-relational data warehouses (PostgreSQL, BigQuery, Snowflake, or Azure SQL). Software Engineering Foundations: Solid grasp of version control (Git), automated testing practices, and API design principles.
Must have: Python, R, Cloud platform, Google Cloud Platform, GCP, Microsoft Azure, Docker, Kubernetes, CI/CD Pipelines, GitHub Actions, GitLab CI, SQL, Data warehouses, PostgreSQL, BigQuery, Snowflake, Azure SQL, Git, Automated testing, API