Expert Exploitation DevOps & SRE (M/F)
Job Description:
WHO ARE WE?
S3NS was born from the industrial partnership between Thales, a global leader in cybersecurity, and Google Cloud, a global leader in cloud solutions. Our ambition is to offer the best of both worlds to all organizations concerned with protecting their sensitive data (public institutions, OIV, OSE.). This means a solution equivalent to Google Cloud Platform (including both GCP IaaS and PaaS services) and respecting the requirements of the SecNumCloud label.
A first offering, 'Local controls with S3NS', has been available since February 2023 to allow our customers to benefit from a first level of transparency and additional controls, and to accelerate the trajectory towards a trusted cloud.
Your Daily Life
As a Compute SRE Engineer, your mission will be at the heart of the operation of our GCP universe. You will be responsible for maintaining operational readiness, the operation, and the management of all Compute platforms (GCE, GKE, Vertex AI.), within an environment equivalent to the GCP universe.You will manage production incidents 24/7, and develop a deep understanding of the technical services that constitute the backbone of the GCP services used by millions of customers.Thanks to our privileged partnership, Google shares its knowledge with us and opens the doors to its internal technical stack (Borg, Colossus, Spanner, Sawmill, etc.) to you. You will benefit from an accelerated and intensive training program, delivered in direct contact with Google experts and our S3NS teams already trained by them, to master exceptional technologies usually reserved for Google's internal engineers.
Key Responsibilities Monitoring SLI/SLO: Monitor the availability, scalability, latency, and efficiency of sovereign GCP services related to COMPUTE by handling production incidents.Infrastructure Management: Manage the machine fleet and the underlying compute stacks (Borg and the various Google control planes) to guarantee optimal performance.Incident Diagnosis and Resolution: Guarantee the maximum reliability of the platform. In order to ensure 24/7 operational readiness of our cloud, the position involves daytime shifts as well as rotating night on-call duties.Team Collaboration: Collaborate with GCP service experts around the world to help mitigate and resolve incidents.Automation & Knowledge: Document knowledge to ensure that all S3NS SREs work with the same information, standardize resolution flows, and improve operational playbooks.***Post-incident Reviews: After an incident, bring teams together to perform a post-mortem, understand the causes, learn lessons, and encourage continuous improvement.
In accordance with the SecNumCloud qualification requirements of our services, delivered by ANSSI, this position is subject to reinforced security requirements. The selected candidate must undergo a security investigation conducted by our services, in accordance with our personnel security policy.
Your Profile
What passions and motivates you: Technological innovation, the Cloud, and the operation of services and infrastructures in "as code" mode. Operational excellence at scale: The operation and high availability (≥ 99.99%) of critical databases and high-volume analytical engines.***The Compute engineering DNA at Google: The curiosity to explore how a hyperscaler manages millions of cores and the desire to master more than 20 years of innovation in distributed systems (GCE, GKE, Borg, Colossus, Spanner, etc.). Impact within a collective of experts: The perspective of joining a specialized team at the heart of Compute challenges (machine fleet, control plane, virtualization, and orchestration).
Experience: Graduate of an engineering school or holder of a Master's degree.At least 3 years of experience in SRE and operations automation. Position open to junior profiles with significant experience in work-study programs in a similar scope.Exposure to an international environment with a good level of English required.***Bonus: Previous experience on GCP is a plus, but your thirst to learn our own Cloud is paramount!
A word from the team
"Want to join a passionate, human-sized team? Join our Compute SRE team: a collective of 7 engineers at the heart of our trusted cloud engine. Your mission is critical: pilot the machine fleet and the underlying compute stacks (Borg, control plane). If you love pure infrastructure and large-scale systems, come and bring your expertise to life with us!"