IT Operations Engineer
Daily tasks
- Incident and problem management: responsibility for the RCA (Root Cause Analysis) process for production incidents: diagnosing problems, resolving them, and implementing preventive actions to ensure they do not recur.
- Monitoring and production environment support: continuous monitoring of service health, early detection of anomalies, and responding before they escalate into incidents.
- Deployment execution: designing, implementing, and maintaining CI/CD pipelines using GitHub Actions and related tools to automate build, testing, security scanning, and application deployment processes.
- Environment oversight: maintaining the stability and consistency of Pre-Production and Production environments. This is not about building them from scratch, but about ensuring their correct daily operation.
- Documentation and knowledge management: documenting operational procedures, known issues, and their resolutions to build a reliable knowledge base for the team.
- Cross-team collaboration: close cooperation with development and platform teams in analyzing and resolving problems, clarifying operational requirements, and ensuring the flow of information between the production environment and development teams.
Requirements
Minimum 5 years of experience in IT operations, application support (2nd and 3rd line support), or a similar role related to maintaining production environments. Experience in end-to-end incident management from the moment of report/alarm, through Root Cause Analysis (RCA), to the implementation of preventive actions. Minimum 2 years of experience working according to the ITIL framework (incident, problem, and change management). Experience working in an Agile environment in collaboration with development teams. Very good knowledge of English (minimum C1), enabling clear communication of technical issues to both IT specialists and non-technical individuals. Very good knowledge of Jenkins in terms of creating and maintaining pipelines and executing and troubleshooting deployment-related issues. Very good knowledge of CI/CD processes, including their automation and continuous improvement. Very good analytical and problem-solving skills, allowing for independent diagnosis of complex production issues. Knowledge of log analysis and alert monitoring tools: Splunk, Sysdig. Knowledge of observability tools: Prometheus, Grafana (reading dashboards, tuning alerts). Experience working with services running on Kubernetes (checking pod status, log analysis, restarting services; without cluster administration). Very good knowledge of Ansible for implementing configuration changes in controlled operational scenarios. Good knowledge of Docker and Docker Compose. Basic scripting skills in Bash and Python to automate repetitive operational tasks and for data agreement/reconciliation.
Must have: Jenkins, CI/CD, Splunk, Sysdig, Prometheus, Grafana, Kubernetes, Ansible, Docker, Docker Compose, Bash, Python