Team Lead - Site Reliability Engineering (all genders)
Must-have:KubernetesCloudDevOpsAIE-CommerceLeadHybrid
Machine translation — original language: German.Show original
Introduction
FACT-Finder develops Product-Discovery technology for eCommerce and is used by leading online shops in Europe with its products Next Generation and Infinity. Currently, we are consistently modernizing our hosting towards Kubernetes on Harvester – as an on-prem hybrid with the option to scale fully into the cloud in the medium term. As Team Lead Site Reliability Engineering (all genders), you are responsible for the reliability, scalability, and costs of our hosting environments, drive this transformation end-to-end, and lead the team that implements it.
Your Tasks
- You are responsible for the operational health of our hosting across On-Premise (Frankfurt, Stockholm) and Cloud – availability, performance, incident management.
- You actively drive modernization towards Kubernetes on Harvester: cluster topology, storage (Longhorn), networking (VLAN, Load Balancing, Ingress), backup, and disaster recovery.
- You build a production-ready k8s platform: lifecycle, upgrades, RBAC, secrets, GitOps (Argo CD / Flux), observability, and policy guardrails.
- You design the NG Search Operator (Custom Kubernetes Operator) and solve auto-scaling (HPA, VPA, KEDA, Cluster Autoscaler) for the current architecture.
- You concretely define our on-prem hybrid model: which workloads run where, how we burst into the cloud, how we keep latency and costs under control – and keep the architecture portable enough for a later cloud-only step.
- You are responsible for capacity planning and hosting costs, making costs a targeted, controllable lever.
- You lead and develop our current 4-person hosting team both technically and disciplinarily, are responsible for performance, and shape the technical standards and ownership culture.
- You make AI a fixed component of our operations: diagnosis, automation, monitoring, and insight.
Your Profile
- Solid background in infrastructure or platform engineering across on-premise and cloud.
- Hands-on depth with Kubernetes in production: cluster lifecycle, upgrades, networking, storage, RBAC, observability, GitOps delivery.
- Proven strong leadership experience, excellent communication, and stakeholder management.
- Ideally, practical experience with Harvester or comparable HCI/virtualization platforms (KubeVirt, vSphere/ESXi, OpenStack).
- Experience with a real migration from bare metal / classic VMs to a k8s-based platform – including stateful workloads, storage migration, cutover, and rollback.
- Confident handling of Kubernetes Operators (Custom Controllers / CRDs), ideally for stateful systems such as search, databases, or streaming.
- Solid understanding of auto-scaling primitives (HPA, VPA, Cluster Autoscaler, KEDA) and their interaction with capacity planning.
- Experience with on-prem hybrid architectures and responsibility for the reliability, capacity, and costs of productive systems.
- Hands-on fluency in using AI tools in operational business.
- Very good English skills; German is an advantage.
THE JOY OF WORKING WITH US
- Impact from day one: Your work directly affects the revenues of leading eCommerce brands in Europe.
- Leadership role with creative freedom: You lead a well-coordinated team and shape our platform in a decisive phase of our transformation.
- Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path towards the cloud – with room to build things new and right.
- AI-first mindset: We do not use AI as a buzzword, but as a fixed part of our daily work.
- Ownership & growth: Clear responsibility, short decision-making paths, and the opportunity to actively shape your role.
- Flexible working: Hybrid working model with a focus on results.
- Strong team: Experienced engineers, an open feedback culture, and an environment in which reliability is taken seriously as an engineering discipline.
- Attractive benefits: Competitive salary, modern equipment, training budget, and regular team events.
Location
Berlin, Munich, Pforzheim or Stockholm (hybrid)