Observability Specialist
The role
If it's running in production, you can see it. You own the stack that turns raw signals into something a human can actually act on, alerts that mean something, dashboards that tell the truth, and coverage that doesn't have blind spots. Done well, this is the difference between an on-call rotation that's sustainable and one that burns people out. Eight of you build and hold that standard together.
What you'll do
Operate and tune the observability platform, owning SLO/SLI instrumentation
Drive alert quality so that every page is actionable, and maintain synthetics and Real User Monitoring (RUM)
Set instrumentation standards for new services
What we're looking for
3+ years of experience in observability or Site Reliability Engineering (SRE)
Hands-on experience with Prometheus, Grafana, and OpenTelemetry
Scripting proficiency
You'll thrive here if you
You treat a noisy or non-actionable alert as a bug to fix, not something to mute and move past
You'd rather instrument a service properly upfront than debug blind later
You believe on-call should be humane, and you build the tooling that makes that true
You measure your work by whether coverage is complete and MTTD keeps improving, not just whether the dashboard looks good