Cloud Operations TM Lead
Accountabilities: Manage and support SaaS Kubernetes applications running in AWS, ensuring reliable and efficient production operations.
Monitor event-driven alerts, investigate incidents, and resolve operational issues to maintain application availability and service-level objectives.
Triage and resolve tickets raised by client services teams, collaborating with relevant technical stakeholders when escalation is required.
Work with solution architects and business stakeholders to identify, design, develop, and deploy infrastructure engineering solutions.
Support the rapid and controlled rollout of new technologies and capabilities across cloud environments.
Troubleshoot operating system, application, virtualization, and containerization issues across the technology stack.
Support large-scale Java applications and contribute to troubleshooting .NET-based applications where required.
Implement and maintain Infrastructure as Code using Terraform and related technologies.
Develop automation and operational tooling using Jenkins, Ansible, Python, Terraform, and similar technologies.
Collect and analyze system and application performance data, identify root causes, and implement appropriate tuning and optimization.
Manage Kubernetes environments, including Helm-templated deployments and containerized applications such as Tomcat.
Perform full-stack troubleshooting across cloud infrastructure, applications, containers, networking, and supporting services.
Use Git and established source-control practices to manage automation and infrastructure code.
Contribute to the continuous improvement of operational processes, automation, reliability, and service-level performance.
Requirements:
Bachelor’s degree or an equivalent combination of education and relevant professional experience.
5+ years of experience in cloud operations, infrastructure engineering, DevOps, or software development environments.
3+ years of experience supporting large-scale Java applications.
3+ years of hands-on experience supporting AWS-based environments.
Extensive practical experience with Infrastructure as Code principles and technologies, particularly Terraform.
Strong understanding of Linux operating systems and cloud-based production environments.
Hands-on experience implementing, managing, and troubleshooting Kubernetes solutions.
Broad automation experience using technologies such as Jenkins, Terraform, Ansible, and Python.
Experience collecting performance metrics, analyzing system behavior, troubleshooting issues, and tuning applications or infrastructure.
Strong knowledge of application containers, particularly Tomcat.
Experience with full-stack troubleshooting and Git/source-control management.
Multi-cloud experience, particularly with AWS and Azure, is advantageous.
Experience supporting .NET applications and working with Dynatrace is a plus.
Familiarity with Helm-based Kubernetes environments and microservices architectures is desirable.
Understanding of common application protocols and messaging technologies, including TCP/IP, HTTP, SOAP, SMTP, REST APIs, XML/JSON, JDBC, and JMS/MQ.
Knowledge of application threading and concurrency concepts, as well as troubleshooting response-time, connectivity, authentication, authorization, and configuration issues.
Strong analytical, problem-solving, and communication skills with the ability to work effectively across engineering, architecture, business, and customer-facing teams.
Ability to work independently in a production-focused environment and continuously improve operational processes.
Benefits:
Full-time, remote position based in India.
Opportunity to work with large-scale SaaS platforms and production cloud environments.
Hands-on exposure to AWS, Kubernetes, Terraform, Jenkins, Ansible, Python, Java, .NET, and modern container technologies.
Broad technical experience across cloud infrastructure, application operations, automation, observability, networking, and troubleshooting.
Opportunity to contribute to cloud modernization and the adoption of new technologies.
Collaboration with solution architects, engineering teams, business stakeholders, and client services teams.
Opportunities to develop automation that improves operational efficiency, reliability, and service levels.
Exposure to multi-cloud environments and geographically distributed production systems.
Inclusive and collaborative working environment focused on innovation, continuous learning, and professional growth.
How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1