Site Reliability Engineer / DevOps
Job Summary We are seeking a DevOps/SRE Engineer with 5–12 years of experience in cloud and platform engineering to design, automate, and support cloud infrastructure using AWS, GCP, or Alibaba Cloud. Responsibilities Develop and maintain cloud infrastructure using AWS, GCP, or Alibaba Cloud to ensure scalable and reliable platform operations Write and optimize Python scripts to automate deployment, monitoring, and maintenance tasks for cloud environments Implement infrastructure as code using Terraform to provision and manage cloud resources efficiently Deploy, manage, and troubleshoot Kubernetes clusters to support containerized applications in production Communicate effectively with cross-functional teams to resolve production issues and improve system reliability Diagnose and troubleshoot platform and application issues to minimize downtime and improve performance Automate repetitive operational tasks to enhance efficiency and reduce manual intervention Provide production support by monitoring system health and responding promptly to incidents Preferred competencies and qualifications Experience with Alibaba Cloud platforms is an advantage