Principal Linux Engineer
Must-have:PythonGitAWSGoogle CloudDockerKubernetesCloudDevOpsCI/CDSecuritySeniorLeadPrincipal
Our client is a leading global financial markets and technology organisation that operates high-availability infrastructure supporting customers around the world.
They are looking for a Principal Linux Engineer to provide hands-on operational ownership of large-scale Linux environments. This is a highly operational, production support-focused role with responsibility for platform stability, incident response, maintenance execution, system reliability, and continuous operational improvement. This position requires weekend support as part of the core working schedule.
Key Responsibilities
- Own the operational health, availability, and performance of Linux infrastructure across production and non-production environments.
- Lead and execute maintenance activities, system upgrades, patching, failover testing, and recovery exercises.
- Act as the senior technical escalation point for Linux-related incidents, driving triage, resolution, and root cause analysis.
- Administer and maintain bare-metal Linux servers, storage systems, and configuration management platforms.
- Execute and monitor production changes in accordance with established change management processes.
- Develop and maintain operational automation using Python, Shell scripting, and infrastructure management tools.
- Support monitoring, observability, and capacity management through dashboards, reporting, and system analysis.
- Create and maintain operational procedures, runbooks, and recovery documentation.
- Collaborate with Engineering, SRE, Security, and Infrastructure teams to improve platform reliability and operational efficiency.
Requirements & Qualifications
- At least 8 years of Linux systems administration and infrastructure operations experience in large-scale, 24x7 environments.
- Deep expertise in Linux OS administration, performance tuning, troubleshooting, and kernel fundamentals.
- Strong experience with configuration management tools such as Salt, Puppet, or Ansible.
- Proficiency in Python and Shell scripting for automation and operational tooling.
- Experience supporting bare-metal server environments and enterprise storage technologies (SAN, NAS, RAID, NVMe).
- Hands-on experience with monitoring and observability platforms such as Grafana and Prometheus.
- Familiarity with cloud and container technologies including AWS, GCP, Docker, and Kubernetes.
- Experience with Infrastructure as Code and DevOps tooling, including Terraform, Git, and CI/CD pipelines.
- Strong problem-solving skills with the ability to perform effectively during high-severity incidents.
Work Schedule
Working schedule: Wednesday to Sunday, 7:00 AM to 4:00 PM (SGT).
Support for early morning operational activities is required to align with global infrastructure operations.