Site Reliability Engineer 3

Jobgether· Brussels (Firmensitz, recherchiert)· lever· objavljeno 31. 07. 2026.
Obavezno:AWSAzureGoogle CloudKubernetesCloudDevOpsAIRemote

Accountabilities: The Site Reliability Engineer 3 will be responsible for ensuring the reliability, scalability, and performance of production systems while implementing modern AIOps capabilities. The role requires strong technical expertise, operational ownership, and the ability to use automation and AI-driven approaches to improve engineering efficiency and customer experience.

Provide end-to-end reliability ownership for production systems, including on-call support, incident response, postmortems, and SLO/SLI management.

Build and maintain observability solutions across metrics, logs, and traces with intelligent monitoring and anomaly detection capabilities.

Design and implement AIOps workflows for event ingestion, enrichment, correlation, deduplication, and alert noise reduction.

Develop AI-assisted operational processes, including automated log analysis, incident summaries, root-cause analysis support, and runbook generation.

Create automation solutions that reduce operational effort through AI-enhanced runbooks, controlled remediation workflows, and self-healing capabilities.

Implement safeguards, approval processes, rollback strategies, and audit controls for automated operational actions.

Own and improve observability platforms, including ELK/OpenSearch environments for logging, indexing, querying, dashboards, and alerting.

Enhance incident response processes through AI-assisted diagnosis, timeline reconstruction, and operational intelligence.

Build and maintain integrations between observability platforms, ticketing systems, ChatOps tools, CMDBs, and automation frameworks.

Partner with software engineering teams to improve deployment reliability, operational readiness, and production stability.

Drive improvements in system resilience, performance, scalability, and infrastructure efficiency.

Maintain high-quality documentation, runbooks, and knowledge resources to support effective operations.

Support capacity planning, performance optimization, and reliability engineering initiatives.

Evaluate AI-generated recommendations critically and ensure appropriate human oversight for high-impact operational decisions.

Measure and improve AIOps effectiveness through operational metrics such as alert reduction, faster incident resolution, and automation adoption.

Requirements:

The ideal candidate is an experienced Site Reliability Engineer with strong cloud infrastructure expertise and hands-on experience applying automation and AI technologies to production environments. You should have a strong problem-solving mindset, excellent operational judgment, and the ability to collaborate across engineering teams.

6+ years of experience in Site Reliability Engineering, AIOps, DevOps, or production engineering roles within large-scale cloud environments.

Strong knowledge of Linux/Unix systems, networking, distributed systems, and cloud platforms such as AWS, Azure, or Google Cloud.

Expert-level experience with ELK/OpenSearch, including log ingestion, indexing, scaling, querying, dashboards, and alerting.

Strong understanding of observability practices across logs, metrics, and distributed tracing.

Experience preparing telemetry data for AIOps implementation, including metadata enrichment, service mapping, normalization, and correlation.

Hands-on experience implementing AIOps solutions such as anomaly detection, alert correlation, intelligent alerting, and automated incident enrichment.

Experience integrating AI capabilities into SRE workflows, including LLM-based incident analysis, operational assistants, or automated troubleshooting processes.

Understanding of prompt engineering concepts for operational use cases, including structured outputs, tool usage, and retrieval-augmented context.

Ability to evaluate AI-generated recommendations critically and identify risks such as hallucinations or incorrect remediation guidance.

Familiarity with agent-based AI frameworks or integrations connecting AI systems with operational tools is a plus.

Strong knowledge of incident management, root-cause analysis, SLOs, reliability practices, and production operations.

Experience with Infrastructure as Code tools such as Terraform, Ansible, or similar technologies.

Ability to work independently, manage priorities, and collaborate effectively in a remote global environment.

Relevant certifications in cloud, DevOps, Kubernetes, AI/ML, AIOps, or observability are preferred.

Benefits:

Competitive compensation package.

Fully remote work opportunity from India.

Opportunity to work with advanced cloud, automation, and AI-driven technologies.

Exposure to large-scale production environments and complex reliability challenges.

Collaboration with globally distributed engineering teams.

Professional growth opportunities through innovative technology initiatives.

Inclusive and collaborative remote-first work culture.

Supportive environment focused on continuous learning, knowledge sharing, and employee development.

Opportunities to contribute to technology solutions that create meaningful real-world impact.

How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best!  Why Apply Through Jobgether? 

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1