DevOps Technical Lead
Role Focus
All positions are Senior / Lead Site Reliability Engineers , focused on enabling engineering teams to build stable, scalable platforms and to perform "boring" (i.e., predictable, low-risk) product releases.
In essence: the SRE function acts as a release and reliability enabler across the organization — ensuring predictable deployments, strong operational standards, and high system resilience.
Key Responsibilities
- Act as embedded SRE within one or more product/engineering teams
- Own service reliability end-to-end: development, deployment, and production
- Build and operate cloud-native platforms based on Kubernetes
- Implement Infrastructure as Code and CI/CD automation
- Develop observability solutions (monitoring, logging, alerting)
- Lead incident response and root cause analysis
- Define and enforce operational guardrails and release frameworks
- Drive continuous improvement in performance, resilience, and cost efficiency
- Enable safe, predictable, and standardized releases across teams
- Coach engineering teams toward higher operational maturity (not just "keep the lights on")
Cloud Hyperscaler
Strong hands-on experience with at least one major hyperscaler (AWS, Azure, or GCP). AWS strongly preferred.
Kubernetes
Hands-on experience: cluster setup, application deployment, stateful and stateless workloads, autoscaling
Infrastructure as Code
Skilled in at least one of: Terraform, Ansible, Chef, Puppet
CI/CD
Experience setting up and managing pipelines with Jenkins and/or GitLab CI
Observability
Proficient implementing monitoring/logging solutions, e.g. Prometheus, Grafana, ELK stack
Programming
Working familiarity with at least one of: Node.js, Golang, Java
Ideal Profile (Experience & Soft Skills)
- Deep hands-on experience in SRE / DevOps / Platform Engineering (please indicate years)
- Proven experience leading or managing engineering teams or technical squads
- Strong ownership mindset; able to operate independently with minimal supervision
- Experience in distributed systems and production-critical environments
- Strong incident management experience, including on-call rotations
- Ability to influence across teams without formal/direct authority
- Prior experience working in embedded team structures (i.e., not a siloed central ops team)
- Strong communication skills in English (spoken and written)
Additional Context
A key focus of this engagement is scaling reliability practices across multiple teams . Candidates are expected not only to operate systems, but to actively:
- Define and evangelize operational standards
- Improve release processes across teams
- Coach engineering teams toward higher operational maturity