Site Reliability Engineer (SRE)
Must-have:Kubernetes
Machine translation — original language: Spanish.Show original
JOB MISSION
Ensure the reliability, availability, and performance of production services through continuous monitoring, automation of operational processes, and efficient incident management, to contribute to the continuous improvement of the technological platform and service stability
ACTIVITIES
- Monitor the availability, performance, and stability of production services to guarantee their continuous operation
- Implement and manage monitoring and observability tools to ensure service visibility
- Define and track SLIs, SLOs, and SLAs to measure and guarantee service levels
- Configure and optimize alerts through monitoring criteria to detect incidents in a timely manner
- Attend to and coordinate the resolution of incidents in production environments through diagnostic, escalation, and follow-up activities to restore service continuity
- Analyze metrics, logs, and operation traces through monitoring tools to identify failures, degradations, and anomalies in services
- Perform root cause analysis and document incidents through post-mortem methodologies to prevent recurrence and strengthen operational reliability
- Implement automation of operational tasks and recovery processes through scripts and standardized procedures to optimize service management
- Execute resilience, load, and stress tests through controlled scenarios to validate service stability and availability
- Monitor and manage infrastructure capacity (CPU, memory, network, storage) by tracking processing, memory, network, and storage resources to guarantee adequate operational performance
- Identify bottlenecks through system behavior analysis to propose performance and reliability improvements
REQUIREMENTS
- Bachelor's/Engineering degree in Systems, Information Technology, Computing, or related fields
- 1 year of experience in production systems operation
- Management of critical environments (high availability)
- Incident management and advanced support
- Experience in Streaming or high-traffic services (Desirable)
- Knowledge in Monitoring and observability (advanced)
- Linux Systems management: performance and failure diagnosis
- Networks: connectivity, latency, load balancing
- Automation: scripting
- Advanced Technical English.
- Requirements- Minimum education: Higher education
- Bachelor's degree 1 year of experience
- Languages: English
- Age: between 25 and 25 years
- Knowledge: Automation, Backend, front end Development, Kubernetes, Process automation, Scientific programming languages
- Keywords: ingeniero, engineers, ingeniera, ing, engineer