Head of Developer Platform & Reliability
Riachtanach:GitKubernetesCloudDevOpsCI/CDSecurityLead
Strategy and Roadmap
- Lead the design, implementation, and maintenance of a comprehensive DevSecOps CI build pipeline that integrates strict automated security guardrails directly into the developer workflow. This includes orchestrating local pre-commit secret detection, software composition analysis (SCA) to detect vulnerable third-party dependencies, static application security testing (SAST) and quality gates to enforce code metrics, rootless container builds, immutable container image scanning, and automated smoke testing before final registry release
- Enforce a strictly decoupled CI/CD and GitOps delivery architecture to maintain a clean boundary between build artifacts and deployment configurations. The candidate will manage and automate the promotion of immutable image tags across Dev, SIT, UAT, and Production environments using a single-branch configuration repo with directory-based overlays to completely eliminate branch-per-environment drift, driving automated reconciliation solely via GitOps pull requests and controllers.
- Own AppSec lifecycle governance, supply chain security, and post-deployment runtime defenses across all environments. Key responsibilities include securing the software supply chain by generating Software Bill of Materials (SBOMs) and signing container images (Cosign), integrating Dynamic Application Security Testing (DAST) via OWASP ZAP to catch runtime exploits in staging, and deploying active cloud-native runtime security cameras (Falco) on production Kubernetes nodes to detect and alert on suspicious container system calls in real-time.
- Architect and maintain enterprise-grade infrastructure observability, log aggregation, and self-healing patterns across all group application clusters. This includes configuring active runtime security instrumentation to detect system-call anomalies on production nodes, establishing metric-driven alerting thresholds, and configuring declarative GitOps controllers (Argo CD) to instantly detect and automatically correct cluster drift back to the designated Git configuration.
- Own the reliability, scale, and disaster recovery strategies for the entire platform delivery lifecycle. The candidate will be responsible for defining and testing deterministic, zero-downtime rollback strategies—such as utilizing automated Git reverts to instantly return environments to previous stable states while managing cluster capacity, isolating failure domains, and ensuring the high availability of critical paths across Dev, SIT, UAT, and Production environments.
Developer Platform Delivery
- Deliver and operate golden paths that a new service can adopt in under a day, with CI/CD, observability, security gates and deployment wired in by default.
- Implement the SDLC as pipeline configuration — branch protection, mandatory review, scanning gates, artifact signing, deployment approvals — so compliance evidence is generated as a build artifact rather than assembled at audit time.
- Execute Security Engineering’s policy-as-code in the pipeline without becoming the arbiter of security thresholds or the help desk for build failures.
- Drive measurable improvement in DORA metrics across the engineering organisation, with a baseline established before the modernisation programme scales.
Reliability Discipline
- Establish and own the SLO framework: SLIs measured where the user is, targets tiered by service criticality, and error budgets that carry a pre-agreed and enforced consequence.
- Own incident management from reliability perspective —severity definitions, blameless postmortems, and closure of postmortem actions.
- Set on-call standards that are humane and sustainable: minimum rotation size, page-load ceilings, compensation, and the expectation that on-call time is spent reducing the causes of paging.
- Run production readiness reviews as a gate for tier-1services entering production, and DR and chaos exercises against contracted RTO/RPO commitments.
- Ensure SRE remains an engineering discipline setting the standard, not a centralised operations team absorbing work that belongs to service owners.
Regulated-Environment Obligations
- Ensure segregation of duties is enforced by the pipeline rather than the org chart: peer review as the approving control, pipeline-held deployment credentials, no standing human write access to production, and time-boxed, session-recorded break-glass.
- Produce audit-ready change and access evidence for HIPAA, HITRUST and PDPA scope as a by-product of normal operation.
- Partner with Security Engineering and Compliance on the SDLC standard, its control mapping and its exception register — without owning the risk-acceptance decision.
Leadership
- Build and lead an engineering team with distinct rhythms, protecting reliability capacity from being consumed by platform delivery work during peak migration periods.
- Ensure architectural authority sits with engineers who operate what they design.
- Influence engineering leaders and product teams who do not report to this role - adoption of the paved road must be earned, not mandated.
- Influence engineering leaders and product teams who do not report to this role - adoption of the paved road must be earned, not mandated.