Senior DevOps Engineer
About Kaseya
Kaseya is the leading provider of AI-powered IT management and cybersecurity software, serving Managed Service Providers (MSPs) and internal IT organizations worldwide. Our comprehensive platform helps organizations efficiently manage, secure, and automate their IT environments, driving operational efficiency and long-term business success.
Backed by Insight Partners , a leading global software investor, Kaseya has experienced sustained double-digit growth and continues to expand its global footprint. Today, Kaseya supports customers in more than 20 countries and manages over 15 million endpoints worldwide.
Founded in 2000, Kaseya has built a culture centered around innovation, accountability, and results. We are a high-growth, high-performance organization that values individuals who are driven, adaptable, and committed to delivering exceptional outcomes for our customers and teammates alike.
At Kaseya, success comes from embracing challenges, moving with urgency, and continuously raising the bar.
About the Role
We are looking for an experienced Senior DevOps Engineer with deep expertise in Linux Administration to join our Backup Platform Engineering team.
In this role, you will own the reliability, scalability, and performance of large-scale Linux infrastructure that powers our next-generation backup and disaster recovery platform.
You will work closely with software engineering, SRE, and platform teams to automate infrastructure, improve operational excellence, and ensure highly available production environments.
Required Skills:
8–12 years of DevOps / SRE experience, with at least 3 years managing large-scale infrastructure
Kubernetes — cluster operations, resource management, custom controllers, multi-tenant workload isolation
Infrastructure as Code — Terraform or Pulumi at production scale; versioned, modular, reusable
CI/CD pipeline ownership — designing and maintaining pipelines (GitHub Actions, Jenkins, ArgoCD or equivalent)
Observability stack — metrics, logs, and traces in production (Prometheus, Grafana, Datadog or equivalent); defining SLOs/SLAs, not just dashboards
Incident management at scale — structured on-call, alert triage, runbooks, post-mortems; experience reducing alert noise (1K+ alerts/month environment)
Networking fundamentals — DNS, load balancing, firewalls, VPC/overlay networks in hybrid environments
Security & compliance mindset — secrets management (Vault), RBAC, image scanning, audit logging; critical for a backup product handling customer data
Scripting proficiency — Go or Python for automation; shell scripting for ops tooling
Linux systems dept h — performance tuning, kernel parameters, storage I/O, process management at scale
Desired Skills:
OpenStack operations — managing Nova, Swift, Neutron, Cinder at scale
Multi-cloud abstraction — managing workloads across AWS, GCP, Azure and private cloud with consistent tooling
Large-scale infrastructure (5K+ nodes) — capacity planning, hardware lifecycle, rack-level failure domains
Cost optimization / FinOps — cloud spend analysis, rightsizing, storage tiering strategies
Chaos engineering — fault injection, game days, resilience testing (Chaos Monkey, Litmus)
Bare metal provisioning — PXE boot, IPMI, automated OS provisioning at scale (Ironic, MaaS)
Backup/DR domain awareness — understanding RPO/RTO, storage replication, data protection pipelines
Go proficiency — reading and debugging Go services, contributing to internal tooling
Additional information Kaseya provides equal employment opportunity to all employees and applicants without regard to race, religion, age, ancestry, gender, sex, sexual orientation, national origin, citizenship status, physical or mental disability, veteran status, marital status, or any other characteristic protected by applicable law.