Senior Site Reliability Engineer
Why this role is special 🧱 Platform-wide impact: You set how reliability, observability and operational readiness work across the whole platform, and every team shipping on it feels the difference.
🤖 Agentic systems in production: You’ll set the guardrails, identity and blast-radius controls for agents operating against real systems and data, and work out what could go wrong before it does.
🛠 Golden paths: You’ll make the platform legible to agents as well as people, through golden paths, MCP servers, skills and documentation that both can actually use.
🚀 Startup pace with scaleup reach: Five years in, we’re working with over 1,000 brands across AU, NZ, USA and the UK, with plenty of scaling still ahead.
What you'll do Build and maintain the cloud infrastructure and paved paths teams ship on, so resilience, security and scalability come as defaults rather than as decisions each team makes on its own
Give teams observability they can self-serve: monitoring, alerting, logging and tracing, plus automation for provisioning, deployments and the operational work nobody should be doing by hand
Make incidents easier to handle: the tooling, runbooks and practice that let whoever is closest to the problem debug it quickly, and post-incident reviews that turn into actual changes
Set the identity and least-privilege model for non-human actors, including agents, CI and MCP servers, so teams can give agents real access without real risk
Document infrastructure designs and operational procedures in a form both people and agents can act on, including machine-readable runbooks
Set the standard for what is safe to ship, and make it easy to meet through guardrails and checks in the pipeline rather than through gatekeeping
Give teams visibility and controls over cloud and inference spend, so the cost of what they run is something they can see and act on
Coach engineers in reliability practices, and work with Engineering and Product to balance feature delivery against platform needs
That's the role, so who are you? You’ve run incident command on real production incidents, and you’ve owned a platform through a meaningful scaling step
Deep infrastructure skills: Infrastructure as Code (Terraform, Terragrunt, CDK), containers (ECS), cloud platforms (AWS preferred, or GCP/Azure), CI/CD, and scripting in Python, TypeScript or Bash.
Strong on observability: Datadog or similar, distributed tracing and structured logging, and a solid grasp of incident management methodology.
Security-minded: Networking and cloud architecture fundamentals, plus an understanding of the security model for agentic systems, including credential handling, least privilege and prompt injection.
Works with agents: You use coding agents in your own work (Claude Code or equivalent) and have a point of view on where they help and where they don’t.
Product-oriented: You navigate technical complexity with business outcomes in mind, and can clearly communicate system health, risks and trade-offs to technical and non-technical people.
Kind: This is a collaborative role, so we ultimately want a great, supportive team player.
Bonus points if you’re passionate about marketing, design, and building exceptional user experiences.
Some of the tools we use: AWS, ECS, Terraform and Terragrunt, GitHub Actions, Datadog, Claude Code, Linear, Notion, Postgres, DynamoDB and Snowflake.