Principal Site Reliability Engineer, Infrastructure & Platform

F5 NETWORKS SINGAPORE PTE LTDSingaporemycareersfuturepublished 09/25/2026
Must-have:Node.jsGitAWSAzureDockerKubernetesCloudDevOpsCI/CDSecurity

Responsibilities: Infrastructure Automation & Configuration Management  Author, maintain, and refactor Ansible playbooks and roles across a large-scale multi-datacenter inventory, covering the full lifecycle from bare-metal provisioning to application deployment  Develop and improve CI/CD pipelines (GitLab CI) for infrastructure automation, including linting, testing, and staged rollout across regions  Manage secrets lifecycle using HashiCorp Vault, including AppRole authentication, secret rotation, and PKI integration  Maintain CMDB/IPAM accuracy in NetBox as a source of truth for all infrastructure assets  Compute & Virtualization  Deploy and manage Proxmox VE hypervisor clusters on bare-metal HPE hardware, including cluster formation, OVS networking, ZFS storage, and VM replication  Provision and lifecycle-manage virtual machines using cloud-init, QCOW2 images, and Proxmox API automation  Manage physical server provisioning end-to-end via HPE iLO (firmware updates, SPP deployment, OS installation via virtual media)  Container & Kubernetes Platforms  Manage self-hosted Kubernetes clusters on-premises, including control plane operations, node provisioning, workload deployment, and upgrade management  Operate Docker-based workloads on infrastructure VMs using compose-driven deployments and container health monitoring  Maintain container image pipelines and registry infrastructure (Azure Container Registry or AWS ECR)  Cloud Platforms  Engineer and maintain infrastructure on AWS and Azure, integrating cloud resources with on-premises systems (DNS, monitoring, identity, networking)  Apply cloud cost awareness, security best practices, and IaC principles (IAM, security groups, networking, storage) across AWS and Azure environments  Networking & Core Services  Operate and troubleshoot core distributed services including authoritative DNS (BIND9), recursive DNS (Unbound), load balancing (HAProxy), and high-availability VIPs (Keepalived/VRRP)  Maintain directory services (OpenLDAP master-replica topology) and AAA infrastructure (FreeRADIUS) used for SSH, VPN, and network device authentication Manage OVS-based network configurations, VLAN topologies, and bonded NIC arrangements across hypervisor fleets 

Observability & Security  Maintain and extend monitoring infrastructure (Prometheus, Observium) across a global fleet including SNMP polling, metrics collection, and alerting  Manage centralised log aggregation pipelines (Fluentbit) and ensure log delivery integrity across DCs  Operate runtime security tooling (Falco) and file integrity monitoring (AIDE) in production environments  Support PCI-DSS compliance activities including CIS hardening, audit logging (auditd), and participation in control reviews 

Reliability & Incident Response  Participate in a 24x7 on-call rotation, responding to and leading production incident resolution Conduct blameless post-mortems and drive remediation of root causes through automation and system improvements  Define and track SLOs/SLIs for critical infrastructure services  Identify and address single points of failure; design and implement HA improvements Requirements: Strong Linux systems administration skills (RHEL/CentOS preferred) including systemd, networking, storage, kernel tuning, and package management 

Proficiency with Ansible (or similar tool) for large-scale configuration management, including role design, inventory management, and CI/CD integration 

Hands-on experience with at least one hypervisor platform, preferably Proxmox VE, Harvester (Kubevirt) or similar (VMware vSphere, KVM) 

Production experience operating on-premise Kubernetes clusters (rke2, k3s, etc) 

Practical AWS or Azure experience including compute, networking (VPC/VNet, security groups, DNS), IAM, and managed services 

Solid understanding of networking fundamentals: VLANs, bonding/LAG, routing, BGP concepts, DNS, load balancing, and firewall rule management 

Experience with secrets management platforms (HashiCorp Vault or equivalent) 

Familiarity with PCI-DSS requirements as they apply to infrastructure -- hardening standards (CIS benchmarks), audit logging, access control 

Experience writing and maintaining CI/CD pipelines (GitLab CI, GitHub Actions, or equivalent) 

Demonstrable on-call experience and comfort leading incident response in a global production environment 

7+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure Engineering role in a production environment