Automation makes most SOC metrics look better, whether or not it helped anyone. Mean time to detect drops because agents pick up alerts the moment they arrive. Alerts closed per hour goes up because closing got faster. Backlog falls. But none of those numbers tells you whether your analysts are better off than they were last quarter. None of them measures a person.
If you asked for automation partly to stop losing analysts, your platform’s metrics can’t tell you whether it worked. The measures that can are mostly in systems you already have. You just aren’t pulling them.
Below are seven measures of analyst workload and strain: what each one is, where to find the data, and how to read it. There are no industry benchmarks here. No credible public ones exist for analyst workload, and we won’t invent them. Every threshold is relative to your own baseline, so the first step is capturing one.
Key Insights: Measuring SOC Analyst Burnout
- Some metrics improve on their own. When agents handle triage, detection and closure times fall even if nobody’s week gets better. That proves the automation ran. It doesn’t prove it helped.
- Averages hide the person who is about to resign. Alerts per analyst is an average. The analyst who quits is often the one carrying twice the median.
- After-hours load is the least measured and most predictive signal. Almost nothing published measures it. And it’s the part of the job people leave over.
- Decision load can rise while alert volume falls. Automation takes the easy cases first, so what’s left needs more judgment. Throughput often improves while strain gets worse.
- Rework is invisible in every standard metric. If an analyst investigates the same pattern four times in a month, that counts as four closures. It looks like productivity.
- Capture the baseline before you automate. You can rebuild most of these from ticket history later. Two you can’t.
Why the Numbers You Already Have Can’t Answer This
Most SOC metrics answer one question: is the operation performing? That’s a fair question, and a lot has been written about it. Our own guide to SOC metrics and KPIs covers the operational and business side in depth. This post is about something else.
The key is knowing which metrics move on their own. When an agent picks up every alert on arrival and closes the routine ones, time to detect and time to first action improve automatically. They’d improve even if the platform did the work badly, as long as it did it fast. So those numbers can’t tell you whether anything improved.
The fix is to pair each automation metric with a measure of analyst workload and read them side by side. That’s the whole method.
| What improves on its own | Read it next to | Why pair them? |
| Time to detect, time to first action | After-hours interruptions per analyst | An agent picks up the 3am alert instantly. Did a person still get woken up? |
| Alerts closed per analyst hour | Decision load per shift | Volume can fall while more of the remaining work needs judgment |
| Backlog size | Queue age, oldest unclosed | A smaller queue full of old cases is a triage problem, not a capacity win |
| Automation rate | Rework rate | Automating work that keeps coming back just repeats it faster |
| Headcount held flat | Workload spread across the team | If headcount stays flat but the spread widens, the benefit landed unevenly |
| Analyst time saved | Median tenure by tier | Time saved only counts if people stay |
The Seven Measures
Workload spread. Compare the busiest analyst’s case count to the team median, per shift, over a month. Pull it from case assignment history. Watch the trend, not the number. If the spread widens after automation, the benefit landed on some people and not others. This is common when automation covers the work most of the team handles, but not one person’s specialty. Say it takes over phishing triage but not the cloud alerts only one analyst covers. Everyone else’s load drops, theirs doesn’t, and the gap widens. The average will look fine the whole time.
After-hours interruption count. How often each analyst gets interrupted outside scheduled hours, per week. Count only interruptions that needed a response, not every alert that fired. Pull it from your paging or on-call tool, not the SIEM. This measure is one of the clearest signals of whether someone leaves, and almost nobody tracks it.
Interruption quality. Of those after-hours interruptions, how many were worth waking someone for? To track it, record a disposition on each page. Most teams don’t do this yet, but it takes an afternoon to set up. A team that gets a few pages, mostly unnecessary, is worse off than a team that gets more pages that were mostly justified. No volume metric will show you that.
Queue age. The median age of open cases, plus the age of the oldest open case, checked weekly. Pull it from the ticket system. Backlog size tells you how much work is waiting. Age tells you how long it’s been sitting, and age is what tracks with the feeling of never being done.
Rework rate. The share of an analyst’s monthly investigations that repeat a pattern they already worked that month. It’s harder to pull but worth it. Group closed cases by alert type and entity, then count repeats per analyst. Standard reports never show this, because each repeat counts as a completed investigation. It’s the measure that tells you when “efficient” really means “repetitive.”
Decision load. The number of judgment calls a person makes per shift. That includes escalations, verdict approvals, containment authorizations and exception decisions. Pull it from audit or approval logs. Expect it to rise as automation matures, since automation takes the easy cases first. Decision load rising while alert volume falls isn’t a problem on its own, but watch it. Forty judgment calls in a shift is heavier than two hundred routine closures.
Tenure and regretted attrition. Median tenure by tier, plus departures split into regretted and unregretted. Pull tenure from HR and get the security team’s view on which departures were regretted. The other six measures are early warnings. Tenure is the outcome they warn about. That’s why it’s on the list, even though it moves last.
Reading Them Against Your Own Baseline
The table below isn’t a benchmark. It’s a way to read your own trend against where your team was three or six months ago.
| Measure | Healthy | Watch | Intervene |
| Workload spread | Busiest analyst near the median | Busiest consistently above the median, same person each month | Busiest at roughly double the median, or the spread widened after automation |
| After-hours interruptions | Falling, and evenly spread across the on-call rotation | Flat while alert volume falls | Concentrated on the same people, or rising |
| Interruption quality | Most after-hours pages justified in hindsight | A noticeable share unjustified | Most pages unjustified, whatever the count |
| Queue age | Oldest open case measured in days | Oldest measured in weeks | Nobody knows the oldest, or cases sit so long no one would act on them anymore |
| Rework rate | Falling as tuning proposals land | Flat while automation rate rises | Rising, meaning the automation is repeating work instead of removing it |
| Decision load | Rising as volume falls, and steady per person | Rising and landing on senior analysts | Rising, concentrating, and escalation quality dropping |
| Median tenure | Stable or rising by tier | Falling at one tier only | Falling at the tier you most need to keep |
Look for the two sets of numbers to disagree. Say your automation metrics are improving, but three or four of these measures sit in Watch or Intervene. That’s the finding this exercise is designed to catch. It usually means the gain went into handling more volume, not into giving anyone their week back. That can be a fair choice. It just isn’t the same as reducing strain, and your numbers should show which one you made.
Capture the Baseline First
You can rebuild five of the seven from systems you already have: workload spread, queue age, rework rate, decision load and tenure. If you’re automating next quarter, you can still get a “before” picture from last quarter.
The other two you can’t. After-hours interruption counts are only as good as your paging records. Interruption quality depends on a disposition for each page, and you can’t add those after the fact. Start tracking both before you automate. They take a few minutes a week.
For a Leader Under Headcount Pressure
Most of these questions come up when headcount is under pressure, so let’s be direct. These seven measures show whether a team is sustainable. They are deliberately not the numbers you’d use to argue for a smaller team.
Throughput numbers show how much work a team can absorb. That leads straight to a conversation about absorbing more. These seven measures show whether the current workload is survivable. If you can show that spread narrowed, after-hours load fell and tenure held, you have evidence the team is sustainable. That’s a stronger position in a budget conversation than a productivity number, which invites the obvious follow-up: could you do this with fewer people?
If the numbers show strain rising while throughput improves, you want to know early. It’s usually fixable by changing which alert types get automated, not by adding people.
Where Conifers Fits
CognitiveSOC™ moves several of these measures because it does the investigation work itself, not just queue triage. It connects threat intelligence, threat hunting, detection engineering, investigation and remediation in one system, draws on each customer’s institutional intelligence, and acts only within limits your team sets.
Two features matter most for analysts. First, every conclusion ships with its reasoning trace and evidence chain, so an analyst reviewing a case reads a structured argument. Second, when evidence is missing, the platform says so instead of reasoning past it. The analyst sees a clearly named gap, not a verdict that only looks clean.
It runs on the security tools an organization already has, with more than 90 integrations. Getting from onboarding to value takes 2 to 4 hours. On its own production data, Conifers measures better than 99% accuracy across roughly 500,000 investigations, with an average investigation time of about four minutes.
Jonathan Waknin’s piece on what gets lost between SOC functions covers the architectural reason repeated context rebuilding happens, which is the mechanism underneath the rework measure above.
Frequently Asked Questions
What metrics measure SOC analyst burnout?
Seven measures cover it: workload spread across the team, after-hours interruptions per analyst, the share of those interruptions that were justified, queue age, rework rate, decision load per shift, and median tenure by tier. None of them shows up in a standard SOC performance report. Most can be pulled from systems you already run.
Why don’t standard SOC metrics show burnout?
Standard metrics measure the operation, not the people. Several of them also improve on their own once agents handle triage. Time to detect falls because alerts get picked up instantly. That shows the automation ran. It says nothing about whether anyone’s week got better.
How do you measure analyst workload properly?
Measure the spread, not the average. Alerts per analyst is an average, and it hides the person carrying twice the median. That person is usually the one most likely to leave. Compare the busiest analyst’s case count to the team median, per shift, and watch whether the gap widens over time.
Can automation make analyst strain worse?
Yes, and the reason is simple. Automation takes the easiest cases first, so the work left behind needs more judgment, even as the total count falls. Alert volume often drops while decision load rises, and every throughput metric will call that a success.
What is rework rate in a SOC?
Rework rate is the share of an analyst’s investigations that repeat a pattern they already investigated in the same period. Standard reports don’t show it, because each repeat counts as a completed investigation. It’s the measure that tells you whether automation is removing work or just repeating it faster.
Should we measure after-hours interruptions or just count pages?
Count the interruptions that needed a response. Then take the step most teams skip: record whether each one was worth waking someone for. A few pages that were mostly unnecessary cost a team more than a larger number that were mostly justified.
What are good benchmark numbers for these metrics?
There are no credible published benchmarks for most of them. If you see a figure presented as an industry standard for analyst workload, trace it to its source before you use it. Compare against your own baseline from three or six months earlier instead. That’s a more honest comparison anyway. Team size, tooling and coverage models vary too much for a cross-industry number to mean much.
How long before automation shows up in these measures?
Queue age moves first. Rework rate and decision load follow once the automated alert types settle in. Tenure moves last, and it matters most. That’s why you track the other six as early warnings instead of waiting for it.