Senior Infrastructure Engineer (Production Platform)
About the Role Within Earth is looking for a highly experienced Senior Infrastructure Engineer to own and manage our mission-critical production infrastructure. Our hotel platform processes over 1 billion searches per day and thousands of bookings daily. This role is responsible for ensuring our production environment remains stable, secure, scalable, and available 24/7 with near-zero downtime. You will be responsible for managing our virtualization platform, Linux and Windows production servers, enterprise databases, load balancing, monitoring, automation, performance optimization, and troubleshooting across the full infrastructure stack. We are looking for someone who can work independently, make sound technical decisions, quickly identify bottlenecks, perform Root Cause Analysis (RCA), and continuously improve the platform.
What You'll Own: Production Infrastructure
XCP-ng Virtualization Cluster
Linux & Windows Production VMs
HAProxy Cluster
SQL Server
PostgreSQL
MongoDB
Redis
ClickHouse
Kubernetes & Docker
Monitoring & Alerting Platform
GitHub CI/CD
Infrastructure Automation
Azure Infrastructure
Production Performance & Capacity Planning
Required Technical Skills: Infrastructure & Virtualization
You should be able to independently:
Deploy, manage, troubleshoot and optimize XCP-ng clusters
Manage VM lifecycle, storage, networking and live migration
Diagnose virtualization performance bottlenecks
Plan infrastructure capacity and resource allocation
Perform disaster recovery and backup management
Linux & Windows Production Servers
You should be able to independently:
Deploy and maintain production Linux and Windows servers
Troubleshoot CPU, memory, storage and networking issues
Perform OS performance tuning
Handle production incidents
Secure and harden production servers
Automate administration tasks
HAProxy
Advanced production experience with:
HAProxy deployment and administration
High Availability clusters
Health Checks
SSL Termination
Sticky Sessions
ACL Rules
Backend Failover
Load Balancing
Zero Downtime deployments
Performance tuning
Troubleshooting production issues
Database Administration
Hands-on production administration of: SQL Server
PostgreSQL
MongoDB
Redis
ClickHouse
You should be able to independently:
Install and configure databases
Backup & Restore
Replication
High Availability
Performance tuning
Query optimization
Capacity planning
Database monitoring
Troubleshoot production issues
Root Cause Analysis
Monitoring & Observability
Experience building and managing enterprise monitoring using: Zabbix
Datadog
Grafana
Prometheus
ELK (preferred)
You should be able to:
Build dashboards
Configure alerts
Monitor infrastructure health
Monitor database performance
Detect anomalies
Investigate incidents
Identify performance bottlenecks
Kubernetes & Docker
Production experience with: Kubernetes
Cluster Administration
Deployments
Services
Ingress
Helm
StatefulSets
Persistent Volumes
Autoscaling
Rolling Updates
Production Troubleshooting
Docker
Docker
Docker Compose
Image Optimization
Container Networking
Private Registry
Production Operations
CI/CD & Automation
Experience with: GitHub
GitHub Actions
CI/CD Pipelines
Release Automation
Rollback Strategy
Secrets Management
Strong scripting skills using:
PowerShell
Bash
Python
Automation experience is essential.
Cloud
Experience managing Azure production environments including:
Virtual Machines
Networking
Storage
Monitoring
Hybrid Infrastructure
Performance & Troubleshooting
We're Looking For Someone Who Can ✅ Own production infrastructure independently. ✅ Troubleshoot across the complete infrastructure stack. ✅ Identify problems before they impact production. ✅ Make sound technical decisions. ✅ Optimize systems before recommending hardware upgrades. ✅ Automate repetitive operational tasks. ✅ Manage high-availability infrastructure. ✅ Design scalable production systems. ✅ Work comfortably in a fast-moving 24/7 production environment.
7+ years in Senior Infrastructure / Platform Engineering.
Experience supporting enterprise production platforms.
Experience managing high-availability environments.
Strong production troubleshooting and RCA skills.
Experience with both on-premises infrastructure and Microsoft Azure.
You should be confident in: Finding production bottlenecks
Performance optimization
Capacity planning
Infrastructure scaling
Database performance tuning
Application infrastructure troubleshooting
Root Cause Analysis (RCA)
High Availability design
Zero Downtime operations