Site Reliability Engineer
Accountabilities: The Site Reliability Engineer will be responsible for enhancing platform resilience, improving incident management processes, and driving proactive reliability initiatives. This role requires strong technical expertise, analytical thinking, and collaboration across multiple teams.
Lead post-incident investigations and perform detailed root cause analyses to identify failures and prevent recurring issues.
Create clear and actionable Root Cause Analysis (RCA) documentation for internal teams and customer delivery.
Develop and implement preventative strategies to improve system reliability and reduce operational disruptions.
Monitor and improve key reliability metrics, including time to resolution and incident response effectiveness.
Configure and maintain observability tools to ensure accurate monitoring, alerting, and performance visibility.
Build client-focused dashboards and alerts to proactively identify application and platform performance challenges.
Collaborate with Engineering, Cloud Operations, and SRE teams to implement improvements that enhance scalability and stability.
Provide guidance and knowledge sharing on effective use of observability tools during investigations.
Create feedback loops with engineering teams to improve development practices, operational patterns, and platform performance.
Contribute to automation initiatives that streamline incident response and operational workflows.
Support SaaS platform reliability by identifying risks, improving processes, and reducing incident impact.
Troubleshoot application and infrastructure issues across cloud environments and enterprise systems.
Requirements:
The ideal candidate brings strong experience in site reliability engineering, cloud operations, incident management, and modern software development practices. They should be comfortable working in complex environments and solving ambiguous technical challenges.
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
5+ years of professional experience in a Site Reliability Engineering role.
Strong understanding of SRE principles, reliability practices, and incident management processes.
Experience with observability platforms such as New Relic, Datadog, Sumo Logic, or similar tools.
Proficiency reading and writing code using technologies such as JavaScript, .NET, and SQL.
Experience troubleshooting C#/.NET web applications and identifying performance issues.
Familiarity with cloud platforms such as AWS or Azure, with AWS experience strongly preferred.
Experience operating within public cloud environments and SaaS platforms.
Strong understanding of cloud architecture patterns and operational best practices.
Knowledge of CI/CD pipelines and Infrastructure as Code (IaC) environments is preferred.
Experience troubleshooting Windows environments and SQL Server is preferred.
Strong analytical, problem-solving, and data-driven decision-making skills.
Ability to work effectively in environments with evolving processes and varying operational maturity.
Excellent written and verbal communication skills with the ability to collaborate across technical and business teams.
Benefits:
Competitive base salary range of $100,000 - $120,000 per year , depending on experience, skills, location, and business needs.
Eligibility for performance-based bonus opportunities.
Medical, dental, and vision coverage for employees and eligible dependents.
Fully paid vision insurance, short-term and long-term disability insurance, and basic life insurance.
Flexible paid time off options and paid company holidays where available.
Hybrid work environment with many positions eligible for fully remote work.
Generous family leave options, including adoption and foster care support.
401(k) retirement plan with company match up to 4%.
Flexible spending accounts, health savings accounts, commuter benefits, and dependent care savings accounts.
Employee Assistance Program providing confidential personal and professional support.
Education assistance programs for certifications and professional development.
Wellness reimbursement programs supporting healthy habits and employee wellbeing.
Additional voluntary benefits, including pet insurance, critical illness coverage, and life insurance options.
How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1