Senior Site Reliability Engineer

WECHAT INTERNATIONAL PTE. LTD.Singaporemycareersfuturepublished 10/01/2026
Must-have:PythonAWSGoogle CloudKubernetesCloudDevOps

Responsibilities 1.Responsible for the operations, monitoring, and resource management of big data solutions on overseas cloud platforms, ensuring the stability and reliability of data platforms and related services. 2.Develop and improve automated operations tools to enhance deployment and maintenance efficiency. 3.Optimize the architecture of data-related services, and design and implement public cloud solutions including fault tolerance and cross-region disaster recovery to improve overall product reliability and quality.

Requirements 1.Minimum 5 years of relevant experience in SRE, DevOps, CloudOps, production operations, or platform engineering. 2.Hands-on experience in resource management, application deployment, and monitoring, with familiarity in technologies such as Kubernetes and container clusters. 3.Familiarity with major public cloud platforms such as GCP and AWS, including the operation and management of relevant cloud services, products, and tools. 4.Familiarity with big data tools and technologies such as Hadoop, Kafka, Flink, Storm, and NoSQL databases. 5.Experience developing automated operations or monitoring tools is preferred. 6.Proficiency in at least one programming or scripting language, such as Shell, Python, or Go. 7.Strong willingness to learn and ability to collaborate effectively within a team.