Senior / Lead Cloud Platform Engineer – Kubernetes & DevOps

Square One Resourcesnofluffjobspublished 09/21/2026
Must-have:PythonAWSGoogle CloudKubernetesCloudDevOpsCI/CDAISeniorLead

Project Overview We are looking for an experienced Senior / Lead Cloud Platform Engineer to join a long-term international project focused on building, evolving and operating modern cloud-native platforms. The role is highly hands-on and combines Platform Engineering, Kubernetes, Cloud Infrastructure, DevOps, SRE and Infrastructure as Code . You will be responsible for designing and improving scalable platform solutions, strengthening automation and GitOps practices, and supporting engineering teams in adopting modern cloud-native standards. You will work in a distributed international environment, collaborating closely with engineering and platform teams. The project offers a high level of technical ownership and the opportunity to influence platform architecture and engineering standards across multiple teams.

Daily tasks

  • Design, build and continuously improve cloud-native platform solutions.
  • Design, develop and maintain Kubernetes-based environments in large-scale production setups.
  • Improve infrastructure automation, provisioning and deployment processes.
  • Design and implement GitOps-based deployment and infrastructure management practices.
  • Build, maintain and optimise CI/CD pipelines and automation workflows.
  • Improve platform monitoring, observability, reliability and operational efficiency.
  • Implement and maintain observability solutions based on technologies such as Prometheus, Grafana and OpenTelemetry.
  • Support and develop cloud infrastructure across AWS and/or GCP.
  • Develop automation and platform tooling using Go and Python.
  • Define and improve Infrastructure as Code practices using Terraform, OpenTofu, Pulumi or similar technologies.
  • Work closely with software engineering teams to establish and improve DevOps, SRE and platform engineering standards.
  • Identify opportunities to improve developer experience, deployment velocity and platform reliability.
  • Take ownership of technical initiatives and drive them from concept through implementation.
  • Contribute to platform architecture, technical decisions and engineering standards.
  • Collaborate with multiple engineering teams and stakeholders in an international environment.

Requirements

8+ years of professional experience in one or more of the following areas: DevOps Platform Engineering Cloud Infrastructure Site Reliability Engineering (SRE)

Strong, hands-on experience with Kubernetes , preferably in large-scale or enterprise production environments. Good programming and scripting skills in Go and Python . Practical experience with GitOps , ideally with Argo CD and/or Flux . Strong experience with Infrastructure as Code , using Terraform, OpenTofu, Pulumi or comparable technologies . Solid experience with CI/CD, automation and deployment pipelines . Hands-on experience with AWS and/or GCP . Strong understanding of observability and monitoring , including technologies such as: Prometheus Grafana OpenTelemetry

Experience designing and operating reliable, scalable cloud-native infrastructure. Ability to work independently and take ownership of complex technical initiatives. Ability to collaborate effectively with multiple engineering teams and influence technical standards. Strong understanding of modern DevOps, Platform Engineering and SRE principles . Nice to Have Experience with any of the following technologies or areas will be considered an advantage: Apache Kafka Temporal OPA / Open Policy Agent Kyverno AIOps and automated remediation AI/ML infrastructure GPU platforms and GPU orchestration Additional open-source cloud-native technologies Experience contributing to or working extensively with open-source projects

Must have: DevOps, Platform Engineering, Cloud Infrastructure, Site reliability engineering, SRE, Kubernetes, Go, Python, Infrastructure as Code, GitOps, Argo CD, Flux, Terraform, OpenTofu, Pulumi, CI/CD, Automation, AWS, GCP, Prometheus, Grafana, OpenTelemetry

Nice to have: Apache Kafka, Temporal, Kyverno, AIOps, Automated remediation, AI/ML