Senior SRE Engineer – Multi-Cloud & Observability

Agile Resources ABGöteborg, Västra Götalands länEURESpublished 09/22/2026
Must-have:PythonJavaAWSGoogle CloudKubernetesCloudDevOpsSenior
Machine translation — original language: Swedish.Show original

Drive observability, reliability, and cost optimization in a complex multi-cloud environment. The role is suitable for you who want to improve platforms, incident management, and production stability.

Do you want to take overall responsibility for reliability, observability, and cost optimization in a complex cloud platform? In this senior SRE role, you will work closely with production-critical systems and drive long-term improvements in a multi-cloud environment.

About the role

  • Take a technical leadership responsibility for SRE, platform stability, and production operations.
  • Design, implement, and further develop observability from infrastructure to business flows.
  • Drive systematic work with reliability, risk reduction, and cloud costs.
  • Participate in on-call duties according to a rotating schedule.

What you will do

  • Build a unified observability architecture for metrics, logs, and traces.
  • Introduce dashboards, alerting levels, SLI/SLO:s, and principles for error budgets.
  • Optimize costs and resource utilization in AWS and GCP according to FinOps principles.
  • Improve Kubernetes clusters, incident processes, root cause analyses, and recovery mechanisms.
  • Develop frameworks for incident management, deployments, rollback, and technical risks.

We are looking for you who have

  • At least five years of experience in DevOps, SRE, or cloud platform.
  • Experience of technical leadership in a senior or leading specialist role.
  • Experience operating large-scale distributed systems in a production environment.
  • Deep knowledge of Kubernetes and its central components.
  • Experience in designing and establishing observability solutions from the ground up.

Technology & Tools

  • AWS and GCP, including platform services, storage, databases, and networking.
  • Kubernetes, EKS, and GKE.
  • Terraform, Pulumi, or CDK.
  • Prometheus, Grafana, Loki, Elastic, Kibana, and OpenTelemetry.
  • Python, Go, or Java; knowledge of eBPF, chaos engineering, and disaster recovery is a merit.

Practical information

  • Location: Göteborg.
  • Way of working: Not specified.
  • Language: English in a professional work environment. Mandarin is a merit.

Contact person

Listed by the employer in the job posting — for questions and your application.