Observability Engineer
Muss:PythonAWSAzureCloudFullstackDevOpsCI/CDSecuritySenior
Observability Strategy and Governance
- Define and own Enterprise Observability Architecture aligned with operational resilience mandates (MAS TRM, DORA, APRA CPS 230).
- Deploy and optimize observability platforms (Datadog, Dynatrace, Splunk) for full-stack visibility across infra, application, network, and user experience.
- Establish governance standards for telemetry data (metrics, logs, traces), ensuring consistency, retention compliance, and security controls.
- Integrate observability platforms with incident management, ITSM, and AIOps systems for predictive alerting and anomaly detection.
Reliability Engineering and Automation
- Implement SRE frameworks for infrastructure and business-critical applications.
- Automate runbooks, alerts, self-healing actions, and auto-remediation workflows via Python, Ansible, and Terraform.
- Partner with Application, Infrastructure, and Cyber teams to codify operational reliability into the delivery lifecycle.
- Conduct resilience testing, chaos engineering, and capacity validation.
- Develop error budget policies and reliability scorecards for key production services.
Cloud Observability and Platform Engineering
- Architect and manage observability for Cloud-native workloads in AWS and Azure.
- Integrate cloud observability into landing zones and CI/CD pipelines for continuous compliance.
- Implement IaC models using Terraform and Ansible for consistent, auditable provisioning.
- Collaborate with Cloud, DevOps, and Security teams on real-time telemetry aligned to audit requirements.
Operational Excellence and Stakeholder Management
- Drive reduction in incident recurrence, MTTR, and manual intervention through observability-led automation.
- Deliver executive dashboards highlighting availability, reliability KPIs, and operational risk indicators.
- Act as technical advisor to senior management during major incidents, post-incident reviews, and audits.
Skillset Requirements:
- At least 5 years of experience in Infrastructure, Cloud, or Site Reliability Engineering (SRE) related roles, with minimum 3 years in an SRE SME capacity, ideally within financial institutions or regulated environments.
- Hands-on expertise with Observability Platforms: Datadog, Dynatrace, Splunk, ELK.
- Hands-on expertise with Automation/IaC: Terraform, Ansible, Python, CI/CD tools.
- Hands-on expertise with Cloud Platforms: AWS (CloudWatch, X-Ray, CloudTrail), Azure (Monitor, Log Analytics, App Insights).
- Deep understanding of SRE principles, service health modelling, error budgets, and auto-remediation design.
- Familiarity with financial sector operational resilience frameworks, regulatory compliance, and incident governance.
- A good team player with excellent written and verbal communication skills, able to coordinate across diverse stakeholders.
Certification in at least one of the following required:
- Datadog Certified Observability Professional / Dynatrace Certified Associate
- Terraform/Ansible/Python Certified Expert
Certification in the following will be advantageous:
- AWS Certified DevOps Engineer / Azure DevOps Expert
- SRE Foundation/Practitioner (DevOps Institute)
- ITIL v4 Managing Professional