Data Engineer Spark Scala Senior | Devoteam Maroc Nearshore

DevoteamRabatJob.bopublished 09/20/2024
Must-have:PythonJavaGitAWSAzureGoogle CloudDockerKubernetesCloudDevOpsDataAgileCI/CDMicroservicesSecuritySeniorLead
Nice-to-have:Fullstack
Machine translation — original language: French.Show original

At Devoteam, we are "Digital Transformakers". Respect, frankness, and passion drive our tribe every day. Together, we help our clients win the Digital battle: from consulting to the implementation of innovative technologies, up to the adoption of usages. Cloud, Cybersecurity, Data, DevOps, Fullstack Dev, Low Code, RPA no longer have any secrets from our tribe! Our 10,000+ collaborators are certified, trained, and supported daily to take on new innovative challenges. A leader in Cloud, Cybersecurity, and Data in EMEA, the Devoteam Group achieved a turnover of 1.036 billion euros in 2022 and aims to double it in the next 5 years. Devoteam Maroc, a reference player in IT expertise for over 30 years (350+ consultants), is accelerating its growth by developing its nearshore expertise activities to meet the needs of our French, European, and Middle Eastern clients. Are you ready to join us and tackle this challenge together?

Data Engineer Spark Scala Senior @ Devoteam Data Driven. In a world where data sources are constantly evolving, Devoteam Data Driven helps its clients transform their data into actionable information and thus make it impactful for more business value. Data Driven addresses the following 3 major dimensions: Data Strategy, Data for Business, and Data Foundation by providing expertise to its clients to make them even more efficient and competitive on a daily basis. Within the Nearshore teams of Devoteam Maroc, you will join the Data Foundation tribe teams: an enthusiastic team of Data Engineers, Data Ops, Tech lead architects, and project managers working on Data platforms and the ecosystem: designing, building, and modernizing Data platforms and solutions, designing data pipelines with an emphasis on agility and DevOps applied to Data. You will be the essential link to provide reliable and valued data to business units, allowing them to create their new products and services, and you will also support Data Science teams by providing them with the necessary "datalab" data environments to successfully carry out their exploratory processes in the development and industrialization of their models, namely: Design, develop, and maintain efficient data pipelines to extract, transform, and load data from different sources to Lakehouse-type data storage systems (datalake, datawarehouse) Write Scala code, often associated with Apache Spark for its concise and expressive features, to perform complex transformations on large volumes of data Rely on the features offered by Apache Spark, such as distributed transformations and actions, to process data at scale quickly and efficiently Identify and resolve performance issues in data pipelines, by optimizing Spark queries, adjusting Spark configuration, and implementing best practices. Collaborate with other teams to integrate data pipelines with SQL, noSQL databases, Kafka streaming, bucket-type file systems … If needed, design and implement real-time data processing pipelines using Spark streaming features Implement security mechanisms to protect sensitive data using authentication, RBAC/ABAC authorization, encryption, and data anonymization features Document code, data pipelines, data schemas, and design decisions to ensure their understanding and maintainability Implement unit and integration tests to ensure code quality and debug potential issues in data pipelines You will reach your full potential through the mastery of your technical fundamentals, your fingertip knowledge of the data you process and manipulate, and above all by asserting your desire to understand the needs and the business for which you will work. Your playground: distribution, energy, finance, industry, health, and transport with plenty of use cases and new Data challenges to tackle together, notably Data in the Cloud. What we expect from you. That you have faith in Data That you help your colleague That you are kind to your HRs That you have fun in your mission And that Codingame does not scare you (you won't be alone: we will help you) And more seriously: That you master the fundamentals of Data: Hadoop, Spark technologies, data pipelines: ingestion, processing, valorization, and data exposure That you wish to invest yourself in new Data paradigms: Cloud, DaaS, SaaS, DataOps, AutoML and that you commit alongside us in this adventure That you like working in agile mode That you create high-performance data pipelines That you maintain this dual Dev & Infra skill set That you are close to the business units, accompanying them in defining their needs, their new products & services: in workshops, defining user stories, and testing through POCs And coding is your passion: you work on your code, you commit in Open Source, you do a bit of competition, so join us What we will bring to you. A manager by your side in all circumstances A Data community where you will find your place: Ideation Lab, Hackathon, Meetup ... A training and certification path via “myDevoteam Academy” on current and upcoming technologies: Databricks, Spark, Azure Data, Elastic.io, Kafka, Snowflake, GCP BigQuery, dbt, Ansible, Docker, k8s … A boost to your expertise in the Data field to become a Tech Lead Cloud (Azure, AWS, GCP …), an architect of future Data platforms, a DataOps expert at the service of business units (Data as a Service) and Data Science (AutoML), a Data Office Manager steering Data Product projects, in short, plenty of new jobs in perspective … The possibility to invest yourself personally: being an internal trainer, community leader, participating in candidate interviews, helping to develop our offers, and why not manage your own team ... A few examples of missions. The design, implementation, and support of data pipelines The deployment of data solutions in an Agile and DevOps approach The development of REST APIs to expose data Support and expertise on Data technologies and deployed solutions: Hadoop, Spark, Kafka, Elasticsearch, Snowflake, BigQuery, Azure, AWS ...

What assets to join the team? Engineering degree or equivalent Expert in the Data field: 3 to 5 years of post-graduation experience Proven mastery and practice of Apache Spark Proven mastery and practice of Scala Practice of Python and pySpark Knowledge and practice of orchestration tools such as Apache Oozie, Apache Airflow, Databricks Jobs Certifications will be a plus, especially in Spark, Databricks, Azure, GCP Mastery of ETL/ELT principles Practice of ETL/ELT tools such as Talend Data Integration, Apache Nifi, dbt is a plus Practice of Kafka and Spark Streaming is also a plus A dual dev (java, scala, python) infra (linux, ansible, k8s) skill set Good knowledge of Rest APIs and microservices Mastery of CI/CD integration tools (Jenkins, Gitlab) and working in agile mode Excellent interpersonal skills, you like working in a team A strong sense of service and commitment in your activities Knowing how to communicate and listen in all circumstances and write without mistakes … and you are fluent in english, indeed!

Additional information. Position based in Morocco in our Rabat and/or Casablanca offices and open only in CDI Hybrid position with possibility of teleworking By joining Devoteam, you will have the opportunity to exchange with your peers, share their experience, and develop your skills by joining the Data Driven community bringing together consultants from the 18 countries of the Group Stay connected: https://www.linkedin.com/company/devoteam https://twitter.com/devoteam https://www.facebook.com/devoteam