Site Reliability Engineer / System Admin
As a Site Reliability Engineer / System Administrator at THCloud.AI, you will be responsible for maintaining the reliability, scalability, and efficiency of our AI and blockchain infrastructure across on-premise and multi-cloud environments. You will drive automation and operational excellence by designing, implementing, and managing CI/CD pipelines, monitoring system health, and proactively addressing potential issues before they impact performance.
Key Responsibilities:
- Maintain, monitor, and troubleshoot the company's cloud, blockchain, AI and associated business systems across on-premise and multi-cloud environments.
- Deploy and manage applications on Linux platforms and virtualized infrastructure (Proxmox, VMware, OpenShift), handling system installations, configurations, and ongoing maintenance tasks.
- Develop, implement, and manage CI/CD pipelines using tools such as GitHub Actions, Ansible, and Kubernetes to ensure seamless and efficient deployment workflows.
- Design high-availability systems with load balancing (HAProxy, Nginx), caching (Redis), and failover configurations.
- Conduct daily monitoring, data backup, and recovery using open-source monitoring tools (Prometheus, Grafana, Loki) for performance reporting, issue tracking, and proactive health checks.
- Perform anomaly detection, root cause analysis, and automated alerting to address and prevent system failures and performance bottlenecks.
- Automate operational tasks and improve system resilience through scripting (Bash, Python, or Golang) and configuration management tools.
- Maintain and optimize infrastructure components such as Docker, Kubernetes, databases (PostgreSQL, MySQL), and distributed storage (Ceph, MinIO).
- Setup VPN, VPC, and secure networking for client environments with proper isolation and security.
- Collaborate with cross-functional teams to support infrastructure improvements, incident response, and operational resilience.
Qualifications: Bachelor’s degree in Computer Science, Information Technology, or a related field, with 4+ years of relevant experience in DevOps, SRE, or similar roles Demonstrated experience with production-grade infrastructure in high-availability (e.g. load balancing) and high-performance environments (e.g. cache optimization) Proficiency in Linux administration and containerization (Docker, Kubernetes) Strong knowledge of CI/CD processes and automation tools (Ansible, Terraform) and experience scripting (Python, Shell) for operational automation Solid understanding of networking protocols (TCP/IP, DNS, DHCP) and networking expertise (VPN, VPC, firewalls)
- Hands-on experience with on-premise virtualization (VMware , ProxMox, OpenShift or similar) and cloud platforms
Proficient in monitoring and logging solutions (Prometheus, Grafana, Loki) for proactive system management Familiarity with database management and distributed storage solutions, particularly PostgreSQL,YugabyteDB, Qdrant and MinIO Multi-cloud and hybrid environment experience Ability to communicate in English at a conversational level
Kontaktinis asmuo
Nurodyta darbdavio skelbime — klausimams ir tavo kandidatavimui.