Senior Site Reliability & System Administrator
This job is no longer listed
The source has removed this listing — applying via the original link is no longer possible.
Must-have:PythonKubernetesCloudDevOpsMicroservicesSecuritySeniorLead
Location: Abuja
Salary: N500,000 - N700,000
We are seeking a Senior Site Reliability & System Administrator with a primary focus on physical infrastructure management, including hands-on hardware setup, configuration, and data center maintenance. In addition to managing, securing, and maintaining our physical and virtual infrastructure, this role owns and enhances application uptime and production performance while bridging the gap between development and operations through robust monitoring, alerting, and incident management.
Key Responsibilities
- Accountable for overall application uptime, reliability, and system performance
- Manage and maintain physical and virtual server infrastructure
- Perform physical server installation, rack mounting, and structured cabling in data center environments
- Lead hands-on hardware provisioning, maintenance, and lifecycle management for servers, switches, and routers
- Ensure physical security of hardware and data center assets
- Design, set up, and configure physical network infrastructure
- Design and manage monitoring and alerting stacks (Prometheus, Grafana, ELK, Datadog)
- Implement and enforce system hardening, security policies, and network administration (VLANs, firewalls, routing)
- Manage backup and disaster recovery processes to ensure business continuity
- Lead incident response, root cause analysis, and post-mortems
- Perform regular system updates, patch management, and hardware audits
- Automate operational tasks and improve deployment processes
- Collaborate with engineering teams to improve service scalability and provide support for infrastructure issues
Required Skills & Experience
- 4–7 years experience in SRE, or system administration roles
- Strong expertise in Linux and/or Windows server administration and networking
- Deep expertise in observability tools (Prometheus, Grafana, ELK, Datadog)
- Strong experience with incident management and on-call rotations
- Strong scripting skills (Bash, Python, Go) for automation
- Experience with Kubernetes and cloud-native environments
- Proficiency in network administration (firewalls, switches, VLANs) and system hardening
- Experience with backup and recovery software and hardware management
- Proven experience in physical data center operations and hardware lifecycle management
- Ability to perform physical hardware installation, cable management, and equipment troubleshooting
- Knowledge of power, cooling, and environmental monitoring for physical server rooms
Nice To Have
- Experience with virtualization technologies (VMware, KVM, Hyper-V)
- Familiarity with configuration management tools (Ansible, Puppet, Chef)
- Familiarity with service mesh architectures and chaos engineering practices
- Knowledge of capacity planning, performance tuning, and distributed microservices
<