Site Reliability Engineer
We are seeking an experienced Site Reliability Engineer (SRE) / Application Support Engineer to support and maintain critical banking applications in a 24x7 production environment. The successful candidate will be responsible for ensuring system stability, resolving incidents, driving service improvements, and enhancing operational efficiency through automation and SRE best practices. Key Responsibilities Provide 24x7 operational support through shift work and on-call arrangements. Monitor production systems and applications to ensure high availability, performance, and reliability. Manage and resolve incidents, problems, service requests, and change activities in accordance with established processes and SLAs. Perform incident troubleshooting, remediation, and escalation to appropriate technical teams where required. Conduct root cause analysis (RCA) for major incidents and implement preventive measures to minimize recurrence. Carry out post-incident reviews and follow up on corrective and preventive actions. Proactively monitor system health, availability, latency, throughput, and end-user experience. Drive continuous service improvements through automation, performance tuning, and reduction of manual operational effort (TOIL). Develop and maintain scripts to improve operational efficiency and support activities across Windows and Linux/Unix environments. Ensure compliance with organizational policies, operational standards, and security requirements. Collaborate with development, infrastructure, database, and business teams to deliver stable and reliable services. Apply Site Reliability Engineering (SRE) principles and best practices to improve system resilience and operational excellence. Requirements Bachelor's Degree in Computer Science, Information Technology, Engineering, or a related discipline. 5 to 9 years of IT experience, including hands-on Application Support in a production environment. Experience supporting business-critical banking or financial services applications. Strong knowledge of Incident, Problem, Change, and Service Management processes. Experience working in a 24x7 support environment with rotational shift duties and on-call support. Hands-on scripting experience in Windows, Linux, or Unix environments (e.g., PowerShell, Shell, Bash, Python). Experience working with relational databases such as Oracle, MariaDB, SQL Server, or similar platforms. Strong troubleshooting and analytical skills with the ability to perform root cause analysis for complex issues. Familiarity with monitoring and observability tools for application and infrastructure support. Understanding of Site Reliability Engineering (SRE) concepts, operational excellence, and automation practices. Ability to manage multiple priorities and respond effectively during critical production incidents. Excellent verbal and written communication skills. Proactive, adaptable, and a strong team player with a continuous learning mindset. Preferred Skills ITIL Foundation or equivalent service management certification. Experience with automation, job scheduling, CI/CD pipelines, or DevOps tools. Knowledge of cloud platforms, container technologies, or observability solutions will be advantageous.