Skip to content
Home » Blog » We Didn’t Fix More Incidents. We Just Saw Them Sooner.

We Didn’t Fix More Incidents. We Just Saw Them Sooner.

The client asked for a 25% improvement in incident detection. We delivered 38%.

More than $2 million in documented savings followed. Here’s the part that still surprises people when I tell the story: we did not get meaningfully better at fixing things. We got better at knowing things were broken.

That distinction is the whole engagement. It’s also, right now, the single largest blind spot I see in businesses deploying AI agents.

The ask was remediation. The problem was visibility.

I was running global infrastructure teams at Kyndryl with 200+ engineers across multiple continents, supporting a Fortune 500 enterprise account. Tens of thousands of endpoints. The kind of environment where nothing is simple and everything is somebody’s dependency.

The mandate coming down was what it always is: reduce incidents, reduce mean time to resolution, reduce cost. Standard. Every operations leader in the world has been handed that sheet of paper.

So we did what most teams do first. We looked at the remediation pipeline. The runbooks, automation candidates, escalation paths, staffing models. All the levers you’re supposed to pull.

The numbers didn’t move the way the math said they should.

The uncomfortable finding

When we pulled the incident data apart, the pattern wasn’t in how long resolution took. It was in how long it took anyone to notice.

A meaningful share of incidents were not being detected by monitoring at all. They were being reported by users. Which means the clock on that incident had already been running. Sometimes for hours before it ever entered a system that measured anything.

Every efficiency gain we were chasing lived inside a window that started too late. We were optimizing the second half of a race we kept entering after the gun.

I’ve watched a lot of organizations make this mistake, including ones I ran. Prevention gets budget because it sounds responsible. Remediation gets budget because it’s visible when it fails. Detection sits between them and gets almost nothing, because nobody gets promoted for finding out about a problem.

What we changed

We moved the investment. Not all of it, and not overnight, but decisively from making the fix faster to making the fault visible.

That meant instrumentation across the estate. It meant fixing the monitoring gaps nobody owned because they sat at the seams between teams. It meant killing alerts that fired constantly and meant nothing, because an alert nobody trusts is worse than no alert because it trains people to look away.

None of this was glamorous. There was no product launch. No architecture diagram anyone wanted to put on a slide. It was months of unsexy, granular work on the least interesting layer of the stack.

It produced a 38% improvement in incident detection against a 25% target.

Why the money followed

The $2 million-plus in documented savings did not come from a cost-cutting program. It came downstream of the detection work, and the sequence matters.

You cannot fix a problem you can’t see. You also can’t price it, staff it, automate it, or make a business case about it. Every efficiency initiative in that environment had been running on estimates, because the underlying data was incomplete.

Once detection improved, three things happened almost mechanically. Incidents got shorter because the clock started earlier. Recurring failures became visible as patterns instead of one-offs, so we could fix causes instead of symptoms. The automation candidates we’d been arguing about for a year suddenly had real volume data attached, which ended the argument.

The savings weren’t the goal we hit. They were the residue of finally having accurate information.

Now put an AI agent in that environment

This is why I keep telling clients that the detection story is the most relevant thing on my resume for the AI era, even though it predates the current wave entirely.

Ask most businesses running an AI agent today a simple question: what did that agent actually do last Tuesday? Not what was it supposed to do. Not what does the prompt say it should do. What did it do? Which systems did it touch? What did it read? What did it change, and who approved any of it?

Most cannot answer. Not because they’re careless, but because they deployed a system that acts on their behalf without deploying anything that watches it act.

I wrote recently about why guardrails written in English aren’t guardrails at all — an instruction in a prompt is a preference, not a control. Detection is the other half of that argument. A control you can’t verify is functionally the same as no control. You’re not secure. You’re just uninformed, which feels similar right up until it doesn’t.

The three-month detection gap in this summer’s AI security disclosures wasn’t a story about model behavior. It was a story about instrumentation, and it’s the same story I lived at Kyndryl with no AI involved.

What I’d actually do first

If you’re deploying agents right now, I’d do this before I’d tune a single prompt.

Write down every system your agent can reach. Not the ones it’s supposed to use but every one its credentials permit. That list is almost always longer than the person who built it expects, and the gap between those two lists is your actual exposure.

Then answer the Tuesday question. If you can’t reconstruct what the agent did on a specific day last week from logs you already have, you don’t have an observability problem to solve later. You have one now, and it’s the reason your first real incident will be discovered by a customer instead of by you.

Then decide what “normal” looks like well enough that abnormal is detectable. This is the hard one, and it’s where most teams stall, because it requires knowing your own process better than the tool does.

None of that requires new AI spend. It requires the discipline of building the foundation before you build on it. The same argument I’ve made about why successful pilots fail in production and what AI-ready infrastructure actually costs.

The thing I’d want a board to hear

Twenty years in enterprise transformation taught me that the organizations getting real returns aren’t the ones with the best models. They’re the ones who can see what’s happening in their own environment.

Detection is not a monitoring line item. It’s the precondition for every other claim you want to make about savings, about risk, about whether the thing you bought is working. Without it, your AI program is a story you’re telling yourself with no way to check it.

I’ve now spent a career on both sides of this. The enterprise side, where I learned it at scale with 200+ engineers and a Fortune 500 environment. The Summit AI side, where I run the same diagnostic for multiple clients who are considerably smaller and, frankly, more exposed because they have fewer people watching and the same agents reaching into the same critical systems.

So here’s the question I’d put to you, and I’d like a real answer, not a comfortable one:

If your AI agent did something wrong last week, would you know yet or would you find out when someone outside your company tells you?

If the honest answer is the second one, that’s not a failure. It’s just information. It’s the most important information you have right now, and it should change what you fund next quarter.


Ready to find out what you can and can’t see? An AI Infrastructure Assessment starts with exactly this question — mapping what your systems actually expose before anything gets deployed on top of them.

Want to see how we work? Our services overview walks through the infrastructure-first sequence, from assessment through implementation and adoption.

More on building the foundation first: The rest of the Summit AI blog covers what breaks, why, and what to do before it does.


Russell Love is the Founder & CEO of Summit AI Business Solutions, based in Browns Summit, NC. With 20+ years of enterprise transformation experience at IBM and Kyndryl, Russell helps businesses build the foundations that make AI actually work.

Leave a Reply

Your email address will not be published. Required fields are marked *