Back to customers

How Recast Software built its entire incident practice on Rootly instead of buying three tools.

Recast Software

x

rootly-logo
"Platform stability is up because the right alerts reach the right people, and the noise never becomes chaos."
Brady LambRecast Software

Brady Lamb

,

Senior Manager of DevOps, SRE, IT, and Security

Recast builds the tools IT departments rely on to manage their endpoints and applications. Its Right Click Tools platform extends Microsoft Intune and Configuration Manager so IT admins work faster and more securely, while Application Workspace manages the full application lifecycle, from packaging to delivery to patching, backed by one of the largest application catalogs in the world.

Founded: 2018, Minneapolis, Minnesota, USA

Size: # ~170

Rootly’s Impact

20%

time savings in manual client communication for maintenance with initial status pages

60%

annual savings in tooling costs

70%

reduction in time to resolution based on infrastructure usage

Almost every fast-growing software company arrives at the same moment: the product works, customers are happy, and then someone realizes the incident management process isn’t keeping up with the growth. Recast hit that moment with monitoring in place but no on-call schedules, so the team was implicitly on call all the time when an alert came in. What Brady Lamb did next is the part worth copying. He went looking for a fix before an outage forced the decision, expecting to buy three separate tools. He ended up with one purpose-built platform.

"We were looking at buying three tools: on-call scheduling, incident management, and status pages. Rootly did the complete lifecycle in one easy to use platform.” -Brady Lamb, Senior Manager of DevOps, SRE, IT, and Security

The stage almost every growing company recognizes.

This is the stage most teams know from the inside. Recast had done the sensible early things: implementing external checks from a down notifier, using the built-in monitoring its cloud services offered, and adding a real dashboarding and monitoring layer. It had also been genuinely lucky, with very few real incidents. What it did not yet have was organization around the alerts. Outputs landed as emails to the whole group, so the team might get pinged about memory pressure without an actual outage alert. And with no defined on-call schedules, everyone carried a pager implicitly, around the clock, without anyone formally being assigned as the point person.

Coordination was improvised the way it is at most companies this size: spin up a channel and a bridge, then track down whoever had the domain knowledge and pull them in. That approach works, right up until it doesn't. As Recast grew and needed to meet SOC 2 Type II and ISO 27001 obligations as well as a 99.9% availability target, the arithmetic changed. More services, more alerts, more people who needed to know what was happening, and no structure to guide it.

So Brady went looking before a painful incident would make the decision for him. He expected to buy three products. Having used other tools in past roles with limited success, what stood out was a platform that could collapse that list into one and still let each team manage its own schedules independently. Rootly was that platform.

One training, one week, one platform.

The rollout matched how a platform team actually works: big bang where it's safe, rolling where it counts. On-call went first, company-wide. After a single training session with Rootly, Recast's managers were building schedules and shifts and showing engineers how to request overrides. Less than a week in, on-call was live for everyone, and the value was immediate.

The new incident response process was rolled out slowly and deliberately, with a second week to define the workflows and services properly (a perfectionist pass, as Brady puts it), and the first product fully live: every alert flowing into Rootly, triggers and workflows firing, status pages up. From there it was a rolling deployment, product by product, with the SRE team continuously tuning which alerts come in and how they route. Next came the customer support organization, on the same model, including API-triggered escalations from systems like Salesforce. Nothing was ripped out along the way as a fail safe.

“On-call was live across engineering in less than a week.” -Brady Lamb, Senior Manager of DevOps, SRE, IT, and Security

What changed

Because Brady owns DevOps, SRE, IT, and security, the payoff shows up differently depending on which of those hats you wear. Here is what changed for each.

For the SRE lead: alerts that route themselves

The technical heart of Recast's setup is the alert pipeline; everything feeds into Rootly, and workflows sort inputs by severity: low and medium alerts flow directly to the responsible engineering team's channel, where that week's on-call engineer triages them, while medium-high and critical alerts become incidents that the SRE and DevOps teams investigate immediately. That sorting is why Brady, when asked whether they deal with fewer incidents now, answers with a firm yes. The alerts that used to involve too many people now land where they belong, and only the ones that matter become incidents. Signal is surfaced without giving everyone access to everything, and every alert that ends in Rootly ties to an actionable item.

For the platform owner: three tools' worth of lifecycle in one

Recast went in expecting to assemble on-call scheduling, incident management, alert coordination, and status pages from separate products. Rootly covered the lifecycle end to end: scheduling with per-team autonomy, incident response with roles and tasks, retrospectives generated from the timeline Rootly logs automatically, and status pages, with no separate systems and no copying files between them. When an escalation point joins mid-incident, they read the description, ask @Rootly to catch them up, and go; nobody has to reconstruct what happened. Even status page updates happen from inside the Slack channel.

For the security and compliance owner: a process you can stand behind

Recast holds SOC 2 Type II and ISO 27001, and runs to a roughly 99.9% availability target. What Rootly adds to that posture is a defined, repeatable process: scheduled coverage instead of implicit 24/7, roles and assignments instead of improvisation, and a complete, automatically logged record of every incident from open to retrospective. The team drills it, too. Rather than waiting for rare but real incidents, Recast runs tabletop exercises against the full workflow, using them to write the playbooks engineers follow, so the process is proven before it's needed.

For the team on call: an AI bot you can just talk to

The newest piece is the one Brady calls awesome: the @Rootly AI agent in Slack. In their tabletop with the first onboarded product, the team updated the incident, created tasks, and closed loops conversationally, "@Rootly, I asked so-and-so to handle this," then "@Rootly, I got this done," and the incident and retrospective updated themselves. No slash-command syntax to remember under pressure.

For leadership: clients who can see for themselves

Before Rootly, Recast had no status pages; client communication was purely outbound. Now status pages are rolling out for every publicly provided service, and clients can check them on their own before filing a ticket or seeing a release, maintenance window, or active issue instead of waiting on support. For a company whose customers are themselves IT teams, that self-serve transparency is exactly the experience they expect, and it converts incidents from support-queue surprises into visible, managed events.

"The AI is what surprised me the most. I tell it in plain language what I want done or ask it for information to help resolve the issue faster. That is the ease of use and the efficiency we have gained from Rootly." -Brady Lamb, Senior Manager of DevOps, SRE, IT, and Security

Most teams don't build an incident practice until growth makes its absence expensive. That's normal, and it isn't a failure. What separates the teams who handle it well is that they start before the outage that would have forced their hand, and they resist assembling the practice out of three or four partial tools. Recast built the whole thing on one platform: live in a week where speed was safe, refined over another where precision mattered, and drilled in tabletops long before it was needed under pressure. If your process today is a Slack channel and a hunt for the person who knows the system, you're not alone. Just know that a better approach to the whole lifecycle is one rollout away.

You and your teams deserve
modern incident management.