Senior Observability Engineer
Vereist:PythonDevOpsAILead
Job Summary
Architect, implement, and optimize Splunk ITSI solutions to enhance enterprise observability. Lead advanced log analytics, anomaly detection, and root cause analysis. Develop automation scripts and integrate ServiceNow CMDB for proactive incident management.
Responsibilities
- Architect and optimize Splunk ITSI solutions including KPIs, glass tables, service analyzers, and service dependency maps to improve monitoring effectiveness
- Lead end-to-end observability across enterprise applications, infrastructure, and distributed systems to ensure comprehensive performance insights
- Perform advanced SPL-based log analytics, event correlation, anomaly detection, and root cause analysis to identify and resolve incidents swiftly
- Configure ITSI Episode Review, correlation searches, and notable event management to enable proactive incident detection and response
- Implement machine learning-based anomaly detection and predictive analytics using Splunk MLTK to forecast and mitigate potential issues
- Integrate ServiceNow CMDB with Splunk ITSI for automated configuration item synchronization and accurate service topology mapping
- Develop Python and Linux automation scripts to streamline monitoring, event management, and operational workflows
- Drive observability best practices aligned with ITIL, SRE, and AIOps frameworks to enhance incident and problem management processes
- Collaborate with stakeholders to analyze performance data, troubleshoot issues, and improve system reliability
Required competencies and certifications
- Bachelor's degree in Computer Science, Engineering, IT, or a related field
- Minimum 10 years of experience in observability engineering, IT operations, or enterprise monitoring
- Advanced hands-on experience with Splunk ITSI, Splunk Enterprise, SPL, and ITSI Service Analyzers
- Proficiency in service modeling, KPI management, glass tables, and event aggregation within Splunk ITSI
- Expertise in ServiceNow CMDB, configuration items, service mapping, and ITSM integrations
- Experience with Splunk MLTK, machine learning, anomaly detection, and predictive analytics
- Strong scripting skills in Python, Bash, and Linux automation environments
- Knowledge of ITIL, SRE, AIOps frameworks, incident management, and problem management
- Proven skills in observability architecture, performance monitoring, and root cause analysis
- Excellent analytical, troubleshooting, and stakeholder management abilities