Site Reliability Specialist, IT Operations

Jobgether· Brussels (Firmensitz, recherchiert)· lever· gepubliceerd op 29-07-2026
Vereist:TypeScriptJavaScriptPythonC#GitAzureDockerKubernetesCloudDevOpsCI/CDAISecurityHybrid

Accountabilities: As a Site Reliability Specialist, you will be responsible for improving the reliability, performance, and scalability of production systems while reducing operational complexity through automation and engineering best practices.

Apply Site Reliability Engineering (SRE) principles to improve the availability, resilience, and performance of production platforms and hosted services.

Design, develop, and maintain automation scripts, workflows, and operational tools to reduce manual processes and improve efficiency.

Build and support production environments while ensuring compliance with operational procedures, security standards, and reliability best practices.

Define and support Service Level Objectives (SLOs), Service Level Indicators (SLIs), and operational reliability standards.

Investigate incidents, perform root cause analysis, and implement both immediate fixes and long-term reliability improvements.

Enhance monitoring, logging, alerting, telemetry, and overall observability to proactively identify and resolve system issues.

Contribute to Infrastructure-as-Code, Configuration-as-Code, Observability-as-Code, and deployment automation initiatives.

Collaborate with development, DevOps, infrastructure, security, architecture, and product teams to ensure operational readiness and continuous improvement.

Explore AI-powered automation tools and intelligent workflows to optimize operations and accelerate incident response.

Create and maintain technical documentation, including SOPs, runbooks, troubleshooting guides, and operational knowledge resources.

Manage incidents, service requests, and change activities while meeting service level agreements.

Participate in scheduled maintenance, platform upgrades, migrations, and an on-call rotation supporting 24/7 operations.

Requirements:

The ideal candidate combines strong infrastructure and automation expertise with a software engineering mindset, enabling continuous improvements across production operations and system reliability.

College or university degree in Computer Science, Information Technology, Software Development, Engineering, or equivalent practical experience.

3–5 years of experience in systems administration, IT operations, DevOps, infrastructure support, automation, or software development.

Experience supporting business-critical production environments and customer-facing platforms.

Strong scripting or programming skills using PowerShell, Python, Bash, JavaScript, TypeScript, C#, or similar languages.

Proven experience designing, developing, testing, documenting, and maintaining automation solutions.

Solid knowledge of Microsoft and/or Linux server administration and production support.

Strong troubleshooting, analytical, and root cause analysis skills.

Good understanding of distributed systems, networking, system performance, reliability, and service operations.

Experience with monitoring, logging, telemetry, observability platforms, and operational analytics.

Familiarity with Git, version control, CI/CD pipelines, deployment automation, and code review practices.

Experience with Infrastructure-as-Code or automation tools such as Terraform, Ansible, Azure DevOps, GitHub Actions, Docker, or Kubernetes is an asset.

Exposure to Azure AI Foundry, Power Automate, AI agents, or similar automation technologies is considered a plus.

Knowledge of cloud platforms, virtualization, backup strategies, and high-availability environments is advantageous.

Strong communication and collaboration skills with the ability to work effectively across technical and non-technical teams.

Excellent English communication skills; French proficiency is an asset.

Industry certifications related to Azure, Linux, Kubernetes, DevOps, or observability platforms are considered beneficial.

Availability to participate in a 24/7 rotational on-call schedule.

Benefits:

Competitive annual salary ranging from $71,330 to $101,900 CAD , based on experience and qualifications.

Hybrid work environment with flexible working arrangements.

Modern collaborative office spaces and access to advanced technologies.

Paid vacation and personal days starting from day one.

Comprehensive employee benefits and savings programs.

Professional development opportunities, ongoing training, and career growth support.

Open communication culture with regular feedback and development planning.

Employee referral bonus program.

Diverse, collaborative, and inclusive workplace culture.

Engaging virtual and in-person social events throughout the year.

Opportunity to work with modern cloud, automation, AI, and Site Reliability Engineering technologies.

How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best!  Why Apply Through Jobgether? 

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1