Senior Data Engineer / Specialist - AWS / Snowflake / Iceberg
Leega is a company focused on efficient and innovative service to its clients. This could not be different with our main fuel: people! Our culture is inspiring and our values are present in our daily lives: ethics and transparency, quality excellence, teamwork, economic, social and environmental responsibility, human relations, and credibility. We seek innovative professionals who are driven by challenges and focused on results. If you are looking for a dynamic and partner company that invests in its employees through constant training, Leega is the place for you! >> LEEGA IS FOR EVERYONE, we will be very happy to have you on our team. Come be part of our history and the construction of our future. Register for our vacancies right now!
Responsibilities and assignments
We are looking for a Senior Data Engineer to work in a lakehouse architecture on AWS, in which Snowflake is the main Data Warehouse and consumes Apache Iceberg tables stored in S3. The professional will be the technical reference for the data ecosystem, responsible for designing, building, and evolving robust pipelines and the interoperability between the data lake and the warehouse, ensuring quality, performance, cost control, and governance of the delivered solutions.
Main responsibilities:
- Design and implement ingestion, transformation, and data availability pipelines in the AWS environment, writing and maintaining Apache Iceberg tables in S3.
- Develop distributed processing jobs in Spark (EMR and/or AWS Glue), including incremental loads, upserts (MERGE), and historical reprocessing.
- Ensure integration between Snowflake and the Iceberg lakehouse: external volumes, catalog integration (Glue Data Catalog), Snowflake-managed tables versus externally managed tables, and metadata synchronization.
- Execute and automate the maintenance of Iceberg tables: small file compaction, snapshot expiration, orphan file removal, schema evolution, and partitioning evolution.
- Model and optimize data models and queries in Snowflake, ensuring performance and efficient use of resources.
- Monitor and optimize cost and performance of AWS and Snowflake environments (warehouse sizing, clustering, cache policies, data lake read cost).
- Act in the definition of data engineering best practices: versioning in Git, CI/CD, data testing, documentation, and lineage.
- Implement governance and security controls: RBAC, masking and row access policies in Snowflake, and permissioning via Lake Formation.
- Diagnose and solve performance and availability problems in critical data environments.
- Act as a technical reference for the team: code review, mentoring less experienced professionals, and recording architecture decisions.
- Collaborate with data, product, and business teams in requirement gathering and solution definition.
Requirements and qualifications
Mandatory requirements:
- Proven experience of at least 5 years working in an AWS Cloud environment.
- Proven experience of at least 3 years with Snowflake, including modeling, query optimization, and warehouse and cost management.
- Proven production experience with Apache Iceberg (ideally 2 years or more): partitioning, schema evolution, snapshots and time travel, MERGE, and maintenance and compaction routines.
- Experience with Apache Spark at scale (PySpark), including job tuning and diagnosis of skew and shuffle.
- Knowledge of AWS services focused on data: S3, Glue (ETL and Data Catalog), EMR, Athena, Lambda, and Step Functions.
- Advanced SQL mastery for complex queries and optimization in large volume data environments.
- Experience with data modeling and Data Warehouse, Data Lake, and Lakehouse architectures.
- Python applied to automation and data engineering.
- Experience with pipeline orchestration (Airflow, Step Functions, dbt, or equivalent).
- Git and versioning culture, testing, and code review.
- Technical English for reading documentation.
Differentials:
- SnowPro Certification (Core or Advanced).
- AWS Certifications (Solutions Architect, Data Engineer, or Data Analytics).
- Experience with Terraform or other Infrastructure as Code tools.
- dbt in production.
- Experience with open catalogs (Glue Data Catalog, Polaris / Open Catalog, Unity) and multi-engine scenarios.
- Experience with streaming (Kinesis, Kafka / MSK, Snowpipe Streaming).
- Observability and data quality tools (dbt tests, Great Expectations, Monte Carlo).
Behavioral competencies:
- Autonomy to conduct technical deliveries with little supervision.
- Clear communication to interact with business areas and multidisciplinary teams.
- Analytical thinking and focus on solving complex problems.
- Collaborative posture and proactivity in proposing technical improvements.
Additional information
Presentation of a PPT case regarding Snowflake performance/usage INDISPENSABLE Knowledge/Experience in:
- AWS (proven experience of at least 5 years).
- Snowflake (proven experience of at least 3 years).
Work model: REMOTE