Cloud Operation Management Engineer
Skal:CloudSecuritySenior
Role Overview
We are seeking an L1 / L2 Cloud Operations & Management Support Engineer to provide day-to-day operational support, maintenance, monitoring, troubleshooting and administration of a Huawei Cloud / Huawei Cloud Stack environment. Depending on experience and capability, the engineer will perform L1 monitoring and initial troubleshooting and/or L2 technical diagnosis and recovery, carry out preventive maintenance, and coordinate with senior engineers or Huawei L3/TAC when vendor-level support is required.
Key Responsibilities
- Provide L1 / L2 operations and maintenance (O&M) support for Huawei Cloud / Huawei Cloud Stack infrastructure.
- Monitor cloud infrastructure health, availability, capacity, alarms and performance.
- Perform daily, weekly and monthly health checks and preventive maintenance.
- Perform L1 initial diagnosis and L2 troubleshooting, as applicable, for incidents involving compute/virtual machines, storage, cloud networking and VPC, virtualisation infrastructure, cloud management platforms, operating systems, backup/restoration, authentication and access-related issues.
- Operate and support Huawei ManageOne and related Huawei Cloud management/O&M platforms.
- Provision and administer cloud resources such as ECS/VMs, VPCs, subnets, storage and associated services.
- Analyse system alarms, logs and performance information to identify faults and service degradation.
- Perform root-cause analysis and recommend corrective and preventive actions for recurring incidents.
- Execute approved configuration changes, patches, firmware/software upgrades and maintenance activities.
- Support backup, restoration, high-availability and disaster-recovery activities.
- Maintain platform security by following access-control, hardening, patching and privileged-account management requirements.
- Maintain technical documentation, configuration records, operational procedures, incident records and maintenance reports.
- Monitor resource capacity and utilisation and highlight potential constraints before service impact.
- Escalate complex or product-related issues to Huawei L3/TAC and coordinate through to resolution.
- Work with Huawei and other infrastructure/application teams during major incidents and planned maintenance.
- Ensure all operational activities are performed according to agreed SLA, incident, problem and change-management procedures.
Minimum Working Experience
- 2–3 years of relevant IT infrastructure, cloud operations or technical support experience.
- At least 2 years of hands-on experience in cloud infrastructure operations, maintenance or technical support.
- Preferably 1–2 years of direct Huawei Cloud / Huawei Cloud Stack operational experience.
- Experience supporting enterprise or production cloud environments.
- Experience performing L1/L2 incident troubleshooting and escalation to senior engineers or vendor/L3 support.
- Practical experience with Linux administration, virtualisation, TCP/IP networking, compute, storage, backup/restoration, monitoring and alarm management, and cloud resource provisioning.
- Experience working in environments governed by formal SLA, incident, problem and change-management processes; ITIL-based operations experience is preferred.