Site Reliability Engineer (FX Production)
Responsibilities:
Respond to production incidents and service disruptions with a structured, data-driven approach to minimise business impact. Identify, investigate, and resolve application and infrastructure issues, provide timely stakeholder communications, conduct root cause analysis, and drive corrective actions, process improvements, and automation initiatives to reduce recurrence.
Maintain the reliability, availability, and performance of business-critical applications and infrastructure. Enhance monitoring and alerting capabilities, implement automation to reduce manual effort, improve operational resilience, perform routine operational checks, monitor system health, and proactively address issues identified through monitoring tools.
Support the planning, implementation, and validation of application releases, infrastructure updates, and maintenance activities. Ensure adherence to change management processes and operational standards, including activities that may occasionally occur outside of standard business hours.
Participate in disaster recovery, business continuity, and resilience testing activities to ensure operational readiness and recovery capabilities.
Collaborate with engineering, project, and business teams to design, implement, and support new systems, enhancements, and operational improvements. Contribute to the delivery of scalable, reliable, and maintainable technology solutions while providing timely support to users and stakeholders.
Adhere to established operational, risk, security, and compliance standards. Support incident, problem, and change management processes while maintaining appropriate controls across technology services.
Share knowledge, support the development of junior team members, and contribute to a culture of operational excellence, innovation, and continuous improvement. Participate in initiatives that enhance processes, tools, and ways of working.
Contribute to broader technology, innovation, learning, and employee engagement initiatives, including workshops, communities of practice, innovation programmes, and cross-functional collaboration activities.
Requirements:
- At least 5 years in production support, SRE, or DevOps in trading or financial services
- Strong Linux/Unix, SQL, and shell scripting, plus Python or Java
- Hands-on with monitoring tools such as Prometheus, Grafana, Splunk, Geneos, or OpenTelemetry
- Good grasp of the trade lifecycle , with FX or other asset class knowledge
- Familiarity with cloud, Docker, Kubernetes, and CI/CD
- Calm under pressure, a clear communicator, and comfortable working independently and with global teams
- Bachelor's degree, preferably in a STEM field