Senior Site Reliability Engineer (all genders)

Omikron Data Quality GmbHBerlin, Munich, Pforzheim, StockholmJob.bopublished 09/03/2026
Must-have:KubernetesCloudDevOpsAIE-CommerceSeniorHybrid
Machine translation — original language: German.Show original

Introduction

FACT-Finder develops Product Discovery technology for eCommerce and is used by leading online shops in Europe with its products Next Generation and Infinity. Both products are currently moving towards a modern, hybrid platform based on Kubernetes and Harvester – with the option to scale fully into the cloud in the medium term. As a Senior Site Reliability Engineer (SRE), you will ensure that our systems remain fast, available, and scalable throughout this transformation. You will work together with the hosting team and experienced engineers and actively help shape the path towards becoming a modern SaaS company.

Your Tasks

  • You define and are responsible for SLOs, SLIs, and Error Budgets across both products and make data-driven decisions regarding reliability and performance.
  • You drive Incident Response: fast detection, clear communication, blameless postmortems, and sustainable follow-ups.
  • You consistently reduce manual work through automation and GitOps (e.g., Argo CD / Flux) and expand self-healing and self-service capabilities.
  • You support the development of an NG Search Operator (Custom Kubernetes Operator / CRDs) and the introduction of Auto-Scaling (HPA, VPA, KEDA, Cluster Autoscaler).
  • You further develop our Observability – Metrics, Logs, Traces, Alerting, and Runbooks that truly help during on-call.
  • You plan capacity and costs across On-Premise (Frankfurt, Stockholm) and Cloud – including burst scenarios into the Public Cloud.
  • You use AI tools to significantly improve diagnosis, alerting, and operational workflows.

Your Profile

  • Experience as an SRE, Infrastructure, or Production Engineer in a SaaS or platform environment – or a strong software/operations background with a clear desire to grow into the SRE role.
  • Solid understanding of SLOs, Error Budgets, Incident Management, and Observability.
  • Hands-on experience with Kubernetes and interest in cluster lifecycle, upgrades, and operator patterns.
  • Experience or strong interest in Harvester or comparable HCI/virtualization platforms (KubeVirt, vSphere/ESXi, OpenStack).
  • Familiarity with GitOps (Argo CD / Flux), Container Storage (Longhorn, Ceph), and Kubernetes Networking (Load Balancing, Ingress).
  • Knowledge of Auto-Scaling primitives (HPA, VPA, Cluster Autoscaler, KEDA) and capacity planning on-prem and in the cloud.
  • Understanding of networks in production-like data centers (including VLAN).
  • Strong automation instinct and an attitude to structurally eliminate Toil.
  • Practical experience in using AI tools in operational operations.
  • Very good English skills; German is an advantage.

THE JOY OF WORKING WITH US

  • Impact from day one: Your work directly affects the revenues of leading eCommerce brands in Europe.
  • Modern Tech-Stack: Kubernetes, Harvester, GitOps, Auto-Scaling, and an exciting path towards the cloud – with room to build things new and right.
  • AI-first Mindset: We do not use AI as a buzzword, but as an integral part of our daily work.
  • Ownership & Growth: Clear responsibility, short decision-making paths, and the opportunity to actively shape your role.
  • Flexible Working: Hybrid working model with a focus on results.
  • Strong Team: Experienced engineers, an open feedback culture, and an environment where reliability is taken seriously as an engineering discipline.

Location

Berlin, Munich, Pforzheim, or Stockholm (Hybrid)