The tempting shortcut
A security team drowning in alerts is the clearest automation case there is, which is exactly why it gets rushed. The instinct is to point AI at the queue immediately and measure the reduction in human review.
That metric will look excellent and mean nothing, because you have automated a triage process that was already producing the wrong answer most of the time.
Three weeks of unglamorous tuning first
In our engagements we spend two to three weeks reducing false positives before automating anything: detection tuning, suppression of known-benign patterns, and correlation that turns four weak signals into one strong case.
Only then do we let agents triage — and their accuracy is measured against the tuned baseline, not the original noise.
What to measure instead
Reduction in alerts reaching humans is a vanity metric on its own. Pair it with investigation depth — what proportion of alerts received a full context review — and with detection recall measured through purple-team exercises. If depth is rising and recall is holding, the automation is genuinely working.
- False positive rate before and after tuning, measured separately from automation
- Investigation depth: proportion receiving full context review
- Detection recall validated by purple-team exercise, not by self-report
- Mean time to contain, split by whether a human was involved
Working on this yourself?
We run ninety-minute working sessions with no charge and no pitch. Bring the process that keeps breaking and whatever numbers you have.
Book a session