Back to customers

How Auror replaced three tools with Rootly and cut the resolution time that the business watches.

Auror

x

rootly-logo
"I've deployed Rootly at two companies now, because I know what good incident management looks like and this is it."
Nigel WrightAuror

Nigel Wright

,

Director of Platform Engineering

Auror is a retail crime intelligence platform with a sharply defined mission: reduce violent retail crime 50% in five years. It captures incidents in a secure, structured way, helping retailers connect the dots to identify organized, repeat offenders and act before harm is done. Used by 85,000 retail stores and over 3,500 law enforcement agencies globally, Auror is working to make stores safer for all.

Founded: 2012 in Auckland, New Zealand

Size: ~300 employees

Rootly’s Impact

3 → 1

tools consolidated

84%

reduction in time to close

0 →1

month full incident visibility

When Auror's platform goes down, the consequence can end in serious harm, not a status page; a dangerous, repeat offender walks into a store with no one warned. So when Nigel Wright joined as Director of Platform Engineering and saw the team was logging about one incident a month, he didn't read it as health. He read it as blindness. The number was low not because the systems were flawless, but because declaring an incident was so painful that engineers avoided it. A quiet channel, on a platform whose whole job is keeping people safe, was the loudest warning in the building. He brought in the incident platform he had already deployed and trusted before.

When silence is the red flag

To most people, having one incident a month seems reasonable. But for someone who had previously been in charge of incident management, that figure was a sign worth looking into. The process back then placed a heavy burden on the person who raised an incident, since the time it took to go from declaring it to closing it was open-ended and could last for weeks, which meant the threshold for classifying something as an incident was higher than anyone had really intended. The teams were still discovering and correcting problems, but this work was not being recorded as incidents and the observability systems had not yet been integrated to pick up on anything that no one had manually reported. Auror was not an unreliable organisation; it just couldn't see how reliable it actually was. In addition, post incident reviews (opportunity to learn moments) were run incredibly rarely (20% of incidents, compared to 95% now), leading to a silent learning loop.

The mechanics underneath were thin. The old stack was Jeli for incident management, PagerDuty for paging, and Freshstatus for the status page, and the workflow was essentially open a Slack channel, then remind someone a retrospective was due. There was no tracking across incidents and no guidance on who should pick up response roles, so during a response one person might be deep in triage while also deciding who to notify and drafting stakeholder comms, switching context constantly in a way that hurt the quality of the response.

"The number our board watches is how fast we respond, and since Rootly that's come down significantly." -Nigel Wright, Director of Platform Engineering

From a bot swap to a lifecycle decision

When Nigel arrived, Auror was already looking for a replacement for the Jeli Slack bot. Four tools were being evaluated and a trial had taken place. At that stage the scope was one-to-one: replacing one constrained Slack bot with another which did about the same job.

In its first week Nigel broadened the question rather than posing it again. Since he had been in charge of Rootly for approximately three years at a previous company, he was able to discuss the capabilities of a complete incident lifecycle platform without the team having to work out the details from the beginning. His point was that Auror should be addressing the entire lifecycle and not just the bot, and it was this reframing that changed a simple swap into a consolidation. As a result, Jeli, PagerDuty and Freshstatus were dropped and Rootly was adopted for On-Call, Incident Response, Retrospectives, the Status Page and Workflows.The most credible endorsement is the one a practitioner makes with their own time, twice.

The rollout split cleanly. Technical implementation was running in less than a month. The people-and-process side took about a month of elapsed time to train every stream, after which a single one-hour session per stream was enough, with an annual refresher since. In the same move, Auror consolidated three tools into one: Jeli, PagerDuty, and Freshstatus all came out, and Rootly went in across on-call, incident response, retrospectives, status page, and all workflows.

The one real risk was paging, and the team named it honestly. PagerDuty is the established name for waking people up. The first big incident, a cloud-related outage in the first couple of months, settled it.

“Rootly’s paging was solid, the escalations fired, and we had all the workflow and post-incident support on top. It does what it says on the tin, and more." -Nigel Wright, Director of Platform Engineering

What changed

Rootly as the routing brain

The feature Nigel calls the game-changer is alert routing. Every alert from Auror's monitoring now flows into Rootly, and Rootly decides what happens; notify in a Slack channel, or wake someone up. Previously that logic lived outside the tool and was hard to visualize. If it pages, Rootly auto-creates the incident so it's ready when the responder wakes. The workflow flexibility extends to calling APIs, crafting request bodies, and validating webhooks, which Auror uses to fire its own custom agents mid-incident to spot commonality across incidents.

Incidents that guide you, with roles that hold

The response is no longer a free-for-all. Workflows assign roles, surface shortcuts and service-specific runbooks, and attach a vendor directory for external dependencies as the front matter of an incident. Background workflows handle the mechanical obligations that used to compete for a responder's attention; flagging when an incident is customer-facing or hits a geography with contractual notice requirements, prompting the team to notify the right customer, and auto-escalating a SEV1 or SEV0 to the SVP of Engineering if it runs past an hour. As Nigel puts it, that's one less thing to think about at the worst possible time. For an enterprise buyer, that is governance and contractual compliance running automatically inside the incident, not living in someone's head.

The post-incident process, finally solved

The blank-page problem is gone. When an incident resolves, Rootly auto-creates a draft of what Auror calls an OTLM (Opportunity To Learn Moment, their term for a blameless retrospective), pre-populated with the timeline, the relevant context, and a reference to what a strong document looks like. Follow-ups and mitigations are tracked, now with SLAs, so the learnings actually get closed out rather than lost.

The number the business watches

In the first year, reported incidents went up, because Auror was finally monitoring and declaring properly. Then it crossed the hump. Time to mitigate is now down 46% and still smoothing, and it's the primary measure Auror reports to its board and executive leadership. That trajectory, worse-then-better as visibility turned into improvement, is exactly what real incident maturity looks like.

Proof points (Auror lightning round, 1-5)

  • Time to first status update improved: 5
  • On-call handoffs are cleaner: 5
  • Stakeholder communications are more consistent: 5
  • More confident handling incidents now: 5
  • Engineering morale is in a better place: 5
  • Pager fatigue is lower: 4
  • Retrospective completion rate improved: 4
  • Customers are happier: 4

A quiet incident channel feels like good news. Often it isn't. It can mean a team has made declaring so costly that people avoid it, which leaves the organization blind to its own reliability. Auror's experienced platform leader recognized that on arrival, brought in the platform he already trusted, consolidated three tools into one, and made incidents safe to declare. The honest result was more incidents on paper first, then a meaningful drop in the MTTR the business now watches, on a platform where reliability is, in the most literal sense, about keeping people safe. If your incident tooling is quiet, that's worth a second look.

You and your teams deserve
modern incident management.