This website uses cookies

Read our Privacy policy and Terms of use for more information.

I’ve spent my entire career in SOCs and worked in all roles there are from analyst to director. In each I was either measured by metrics or responsible for producing them.

Safe to say I’m responsible for a metric (hah) ton of SLA reports and yet, somehow it always felt like those gave no valuable insights into how our SOCs actually operated.

In this blog post I’m highlighting issues I have with how we measure the performance of SecOps, in a follow up article I’ll share my ideas on how we can do it better.

This is part of my healing process so bear with me through this lengthy rant.

Everybody understands metrics differently

I encourage you to do a little field research. Go ask 10 people what MTTR is and how they measure it. Chances are you'll get 10 different answers.

For some the “R” in MTTR stands for resolve. That mostly means a full end-to-end handling of the alert. Sure, that’s actually how I understand it as well.

For others it’s MTT-Respond. But how do we understand response? For some it’s the initial acknowledgement of the alert. That’s also often measured as MTTA or Mean Time To Acknowledge. For others, response ends at an alert being escalated to a higher tier or at an initial notification to the client.

Then, we have a group who claims that “R” in MTTR stands for anything from “Repair”, “Restore” to “Remediate” or “Report”. All answers I got from SOC leads I talked to in the recent months.

Also, it’s no wonder that different interpretations of MTTR exist because MTTR is genuinely different for MSSPs and in-house SOCs. While in-house SOCs mostly measure it as the full time it took to close an alert, MSSPs are often responsible for the initial part of the investigation (triage + maybe some L2 work) and will measure MTTR only for their portion of alert handling.

This means that MSSPs will often see shorter MTTR and higher alert volumes than in-house SOCs even though they might have a common understanding of what MTTR is.

This is just one example, but it demonstrates how a single metric can have multiple different meanings, depending on who you ask. This is why when I hear “our MTTR is 30 minutes” this tells me absolutely nothing.

Overemphasizing speed and volume

Sometimes it’s not even if we understand those metrics the same way, but if those metrics are genuinely useful for us. After all, what you measure is what you optimize for.

And what do SOCs often measure? Speed and volume.

So what do those SOCs optimize for? Speed and volume.

Doesn’t this encourage surface level investigations?

On one hand, we say we want quality investigations. On the other hand, we judge analysts by how fast they handle alerts or how many they can fit into their shift.

You see how that makes no sense?

Some would argue that it’s smart for MSSPs to emphasize speed over quality if they handle the initial triage, and sure, maybe that’s the case. In my experience that can turn into becoming a mass escalation engine, and I know that clients absolutely hate that.

And look, I get it. For a non-technical stakeholder who is used to the “traditional” SOC metrics, having a short MTT-whatever is probably of paramount importance. But I believe this is a slippery slope. Short SLA commitments are often what “sells” your service to a client, but also makes your analysts’ lives that much harder. And if you keep adding clients and all of them need their alerts handled within a very short window, you then have an army of analysts trained to look for a first sign an alert is a false positive. In other words, you groomed your SOC into being mediocre.

Oh, and don’t get me started on alert volume. Those mean nothing in isolation. As in, what is a good number of alerts?

  • If the number is high, does it mean that we did a good job investigating, or that we’ve done a terrible job at finetuning?

  • And what if it’s low? Does that mean that we’ve finetuned our environment perfectly or that we have visibility gaps?

Alert volume is probably the most frequently cited number when trying to demonstrate that SOCs are doing poorly (FUD marketing doing well as ever). It means that you see claims like “SOCs are facing 10,000 alerts / day” in every other post from a cyber vendor and that is often parroted by the less experienced practitioners.

What’s genuinely upsetting is that each of those outlandish volume claims can be supported by an official sounding report. Pick a number of alerts you think a SOC is handling each day, then look for a report justifying it. Chances are, you will find it:

Company / Sponsor

Report (year)

Alert volume claim

Source

Anvilogic / ESG

Trends in Modern Security Operations (2022)

~286 alerts/day per SOC mean; ~99% false positive rate; 96% tradeoff efficacy/efficiency

Cisco

Security Capabilities Benchmark Study (2017)

44% of SOC managers see >5,000 alerts/day;

Crogl / Ponemon Institute

State of SecOps and AI in the SOC (2026)

4,330 alerts/day; only 37% investigated; 16 cyberattacks/year avg

Cybereason

"Eliminate Alert Fatigue" whitepaper

11,000 alerts/day; 45 tools avg; ~50% false positive; 30% ignored

Forrester (commissioned by Palo Alto Networks)

2020 State of Security Operations

Over 11,000 alerts/day average; 70% of analyst time on triage/response

Prophet Security

State of AI in the SOC (2025)

960 alerts/day avg; 3,000+ for large enterprises; 40% uninvestigated

Trellix

Elevating the SOC Analyst Experience With XDR (2023)

Typical company: 10K daily alerts;

Vectra AI

State of Threat Detection (2023)

4,484 alerts/day; 67% ignored

Vectra AI

Defenders' Dilemma (2024)

3,832 alerts/day; 62% ignored

Fortinet

Information Overload: Making Sense of Security Data

Between 10,000 to 150,000 alerts per day

There are many issues with the methodology of those reports, probably the biggest one is that when talking about alert volume you absolutely need to discuss team size. 400 alerts per day for 2-person SOC and 20-person SOC is entirely different workload. Add to it how some of those numbers are taken pre-finetuning, mixing up MSSP and in-house SOC numbers and a real chance that some of those numbers are made up and you’ll see why I’m super skeptical when a vendor says “SOCs are facing 10,000 alerts / day”. Sure.

Incident Response metrics mixed with SOC metrics

One thing I’m noticing recently is that many SOC leaders mix up Incident Response metrics with regular SOC metrics. Just this week I’ve seen two comments about MTT-Contain in the context of Security Operations. H…how much are you containing to derive a mean time from it?

SANS IR framework

What’s even more curious is MTT-Detect. To me this is an incident response metric that describes the time from the initial compromise to its detection. Otherwise what are we measuring? How fast our tools work?

Think about it, we don’t detect anything ourselves. We configure the tools that do so. Our ability to influence the time it takes our tools to detect is limited. We can maybe make detection rules run in a shorter interval, optimize data pipelines, account for ingestion latency, sure. But calculating the time from the original event to the alert firing for every alert (mean) and using it to demonstrate your SOCs success? I’m not buying it.

Oh, and I’ve recently seen SOC leads talk about MTT-Eradicate so there’s that. Again, just how much are you eradicating, good sir?

The curious case of Benign Positives

“True Positive. False Positive. True Negative. False Negative. Long ago, the four metrics lived together in harmony. Then everything changed when the Benign Positives attacked.”

If you’ve worked in a SOC (I assume you have since you’re reading this) you probably saw some type of this sad table that explained how alerts are categorized:

Cool. Actual attack? True Positive. Alert fired for some random activity? False Positive. See how easy that is?

Or rather was, because vendors started adding their own alert resolutions, like for example Benign Positive. And sure, I understand why. Sometimes an alert fires for an activity that looks malicious, but upon investigation ends up being benign.

On one hand we probably needed a way to categorize alerts that require investigation but aren’t malicious to separate them from clear False Positives where you can just close them after a quick look around. On another, that ends up being confusing.

Why you ask?

Consider two SOCs. SOC 1 has a SIEM that uses Benign Positive alert category. SOC 2 operates only on True Positive and False Positive statuses.

Both SOCs received 100 alerts, 50 were clear False Positives, 50 looked malicious and required further investigation, but in the end only 5 were malicious and ended up being raise as security incidents.

For SOC 1 their TP rate is 5%, BP rate is 45% and FP rate is 50%.

For SOC 2 their TP rate is 5%, FP rate 95%.

Both SOCs are in the same situation - same workload, 100 alerts, 5 security incidents, but SOC 1 reports a 50% FP rate, while SOC 2 reports a 95% FP rate. Very different numbers for the exact same outcome.

Wrapping up

So that's my list of grievances. Metrics nobody defines the same way, metrics that optimize for the wrong behavior, metrics borrowed from incident response, and resolution categories that make comparing two SOCs impossible. Am I saying we should stop measuring SecOps?

Absolutely not.

In the next article I'll share what those are, based on what worked (and what didn't) in the SOCs I've run or consulted for.

Reply

Avatar

or to participate

Keep Reading