Production Support Engineer, Service Recovery (Banking, 1-year renewable contract)

EVOLUTION RECRUITMENT SOLUTIONS PTE. LTD.Singaporemycareersfuturepublished 10/08/2026
Must-have:PythonAWSAzureCloudFinTechSecurity

Dear Applicant, If you or someone you know is interested, please send the CV directly to quynh.nguyen@evolutionjobs.sg (most preferred, as I may overlook some CVs due to the high volume). Please note that visa sponsorship is not available at this time. Key Responsibilities Provide end-to-end 24x7 production support for customer-facing banking applications, transaction-processing services and underlying technology infrastructure. Investigate and resolve customer-impacting issues including payment failures, transaction delays/reversals, reconciliation exceptions, login/authentication issues, duplicate or rejected transactions, and card-related incidents. Support production incidents across core banking, retail and wholesale banking, cash management, customer information, credit/debit cards and payment services. Perform end-to-end transaction tracing across applications, databases, APIs, middleware, operating systems, networks, storage and downstream services. Analyse application/system logs, SQL data, batch records, monitoring alerts, error messages and file-transfer results to identify root causes. Troubleshoot issues across application, database, scheduler, operating system, network, storage, middleware, external interfaces and third-party services. Support Start-of-Day/End-of-Day processing, production releases, application deployments, service restoration and Disaster Recovery exercises. Perform or coordinate infrastructure activities including network troubleshooting, OS/storage patching, vulnerability remediation, server deployment, migration, upgrades and capacity/performance checks. Support physical and virtual server environments, infrastructure provisioning, environment builds, lifecycle management, high availability, backup and recovery. Safely perform application, middleware, database and batch service restarts during maintenance, including pre-checks, post-change validation and health checks. Monitor systems, transactions, batch jobs, interfaces and file transfers using tools such as Splunk, Geneos, Control-M and SQL monitoring platforms. Develop and maintain dashboards, alerts and operational monitoring to improve service reliability and visibility. Automate operational and infrastructure activities using Ansible, UNIX/Linux shell scripting, PowerShell, Python and SQL. Apply Incident, Problem and Change Management practices and participate in major incident recovery to restore services within agreed SLA/KPI targets. Conduct root-cause analysis and post-incident reviews, driving corrective and preventive actions for recurring issues. Prepare change records, risk/impact assessments, implementation and validation plans, rollback procedures and production-readiness documentation. Coordinate with business users, banking operations, application, database, infrastructure, network, cybersecurity, command-centre and vendor teams. Maintain SOPs, runbooks, recovery procedures, implementation plans, incident reports and knowledge articles. Ensure all production activities comply with technology risk, information security, regulatory, audit, access-control and change-governance requirements. Key Requirements Bachelor's degree in Computer Science, Information Technology, Engineering or a related discipline, or equivalent relevant experience. 5+ years of hands-on experience in banking application support, infrastructure operations, production support or a combination of these areas. Mandatory experience in banking or financial services, supporting mission-critical systems in an SLA-driven environment. Proven experience handling customer-impacting and high-severity production incidents in a 24x7 environment. Strong hands-on experience investigating payments, transactions, authentication, core banking, cash management, card processing, batch, interfaces and file-transfer issues. Strong capability in end-to-end transaction tracing, log analysis, SQL investigation, data validation and root-cause analysis. Hands-on experience with multiple enterprise platforms such as UNIX/Linux, Red Hat Enterprise Linux, IBM AIX, Windows Server, AS/400 and/or mainframe. Experience with Splunk, Geneos, Control-M, SQL monitoring and enterprise application/infrastructure monitoring tools. Good understanding of networking, OS/storage patching, server deployment, VMware or equivalent virtualization, high availability, backup and Disaster Recovery. Practical automation/scripting experience using Ansible, UNIX/Linux shell scripting, PowerShell, Python or SQL. Strong knowledge of Incident, Problem and Change Management, including major incident recovery, implementation planning, risk assessment, validation, rollback and post-incident reviews. Strong analytical, troubleshooting, documentation and problem-solving skills. Excellent written and verbal communication and stakeholder management skills. Ability to work independently and under strict time constraints, with strong focus on customer impact, transaction integrity, operational risk and service availability. Willingness to participate in rotating shifts, 24x7 on-call coverage, weekend/public-holiday support and after-hours production releases or maintenance. Relevant certifications such as ITIL, RHCSA/RHCE, IBM AIX, Microsoft, VMware, Ansible, cloud, database or infrastructure certifications are advantageous. Experience with retail/wholesale banking, payments, reconciliation, clearing/settlement, APIs, middleware, managed file transfer or Oracle/Microsoft SQL Server is an advantage. Exposure to Azure/AWS, cybersecurity, vulnerability management, regulatory compliance and audit controls is an advantage.

Contact person

Listed by the employer in the job posting — for questions and your application.