AiOps & Observability Specialist
Responsibilities:
Implement and maintain monitoring solutions to track system health, availability, performance, and reliability.
Monitor system logs, metrics, alerts, and dashboards to identify performance issues and potential failures.
Detect, investigate, troubleshoot, and resolve system incidents promptly.
Perform root cause analysis on recurring incidents and implement measures to prevent future occurrences.
Develop and maintain automation scripts to streamline operational processes and reduce manual tasks.
Support proactive performance management by identifying system bottlenecks, capacity issues, and potential risks.
Manage and monitor cloud infrastructure and services across relevant cloud platforms.
Configure and maintain monitoring, alerting, logging, and observability tools.
Troubleshoot application, infrastructure, network, and system-related issues.
Collaborate with development and IT teams to improve system reliability, scalability, and performance.
Participate in incident response, escalation, and post-incident reviews.
Maintain documentation of system configurations, operational procedures, incidents, and troubleshooting processes.
Continuously identify opportunities to improve system performance, availability, and operational efficiency.
Ensure systems and operational processes follow established security, reliability, and compliance standards.
Requirements:
A minimum of 5 years of experience A minimum of a degree in a related field
Top of Form
Bottom of Form