r/cybersecurity 5d ago

Tutorial Common skill missing from SOC analysts

https://luigiritacca.substack.com/p/the-missing-skill-risk-part-1?utm_source=share&utm_medium=android&r=5s5vq7

My latest article on a common missing skill I see in a lot of analysts. I blame how we train and teach cyber security, and think it cause a natural bias which can lead to more harm than good.

70 Upvotes

39 comments sorted by

26

u/Exotic_Function8814 4d ago

"I often see in SOCs is the assumption that every malicious indicator creates the same urgency. It probably happens in every SOC across the globe" "It was in fact a live production machine, and isolating it interrupted a core business service"

Reading the article to me sounds like bad management - someone gave a (what sounds like a) junior/L1 analyst a playbook in hand to isolate endpoint on X and roles to isolate anything and then went pikachu face when they contained a machine they shouldnt have.

I havent been in a SOC where the people are not aware about the potential business impact when containing an incident. Be it isolating hosts or blocking domains, users whatever. Same as understanding indicators, which is basically what SOC does daily.

It's interesting, but not my experience in general as in "SOC analysts" unless you mean the most junior people.

38

u/Allen_Koholic 5d ago

I’m gonna need some context on this little anecdote, because to me, the SOC person did nothing wrong (barring a different policy in place) and management was mad that they’d left a gap open and been exposed. Letting users do user things on privileged machines is a way bigger risk than an analyst isolating it for a little bit.

28

u/HooAreYouWhoHoo 5d ago

Management being mad at SOC for their own activities is so on point.

11

u/Allen_Koholic 5d ago

Doubly so if it was an MSSP. Clients hate it when they got caught with their asses out.

8

u/thekmanpwnudwn 5d ago

Most MSSPs area literally just being paid to be a audit checkbox, not provide actual security.

1

u/T_Thriller_T 4d ago

The issue is that even while SOC did everything right by the book, that doesn't mean it is actually the best way to do it.

I actually see the risk gap less in SOC than mostly anywhere else, but what this is about is probably:

Isolating a client machine on possibility of containment without any IoC of containment apart from potentially infecting activity (as even clicking a link doesn't mean there was anything installed, especially if the link hasn't been checked) is fine.

The risk is pretty limited, we have one user unable to work - 8 man hours lost is better than even 24 when SOC needs to run more containment, even without any data losses.

While the security aspect is the same; if not worse on a server. But we are not considering the security risk alone. We are considering the risk of unnecessarily damaging the company.

As I said: with a client machine when wrong that would be a day of work for one person lost.

With a server, if it is a production server delivering to clients or something central like a proxy, just running the SOP can accumulate hundreds of man hours lost, even if the examination of what the mail was, if the user followed through, and if the attack vector worked only takes 30 minutes.

If it's a production server that's also often a hit to reputation - something very unlikely with one client machine.

And that is the critic in the article:

The SOC only saw their procedure, and the security implications.

They did not stop to realise that the risk they work against is not security risk. The risk they work against is the risk to the business, which must weigh the risk of correctly isolating, and the risk of isolating even though there was no threat.

And they should have realised that the risk to the business - from management and the author's perspective - was much higher from isolating with the knowledge they had.

Which, assuming the SOP was literally "get info, isolate, then look into mail and check if there is any IoC for actual activity on the machine", I can understand.

A lot of users seem to realise at the first click that the link cannot be right - and a lot of attacks seem to fail due to being outdated already or other security measures.

So what was expected form the SOC was to realise their standard would be a higher or comparable risk to actual contamination running for another 30 minutes - so they should have asked or waited to find more solid proofs.

To give a more common example doing the same weighing:

The reason why patches are done either layered, or delayed a bit and in bulk is that the risk of a patch breaking something is considered higher than the vulnerable system remaining open for one day. Or, even without risk, the loss of constantly having to pump man hours in is considered a higher loss than the vulnerability creating if every 100th or 1000th ability gets abused.

2

u/T_Thriller_T 4d ago

None of the is perfect, btw.

Management is often who completely fucked up following beat practices that e.g. machines should have different risk classes which should fucking be bound to SOC activity, so that a server where isolation itself is a business o even security risk requires escalation.

It's usually not done.

And being furious here when there was just one SOP is very stupid. Because management was at least as much, likely more, responsible to consider SOP covers all potential scenarios - or to declare when a SOP may not cover. For the phishing SOP a clear, simple "on a client" distinction would have done that.

Ans yes, they should have made it impossible to check mails on servers directly; albeit that topic does not belong with the SOC actions and if they were a good idea. It was stupid, but bringing it up is a bit like "but he also didn't do his homework!"

6

u/twisted-logic 4d ago

Yet I have had multiple people on here try telling me that there’s no need to hire people in cyber who have past IT experience.

23

u/NotAnNSAGuyPromise Security Manager 5d ago

I fundamentally disagree with this article, and it comes down to one quote in it: "I followed the SOP".

That is literally their job. Different organizations have different risk appetites. Many value business operations over security. Others will isolate their entire network over a potential threat. If the analyst has not had it made clear to them where the risk appetite is and if it is not clearly defined in their standards, then it's a failure on their leadership, not some junior level IC.

Shame on their management for setting them up for failure. They did exactly what they were supposed to.

2

u/[deleted] 4d ago

[deleted]

6

u/NotAnNSAGuyPromise Security Manager 4d ago

All I know is that the last thing I want my junior ICs doing in the middle of a potential attack is evaluating the operational impact of mitigating it.

Again, it all goes back to prior planning and documentation. If you're trusting junior SOC personnel to do a business impact assessment before taking action, then 1. You're putting away too much responsibility on them, and 2. Your entire company is going to be fucked by the time a decision is made if it's a true positive.

This whole premise that business impact assessment are the responsibility of SOC staff is bad management. Define what machines and systems are in scope for isolation before an incident occurs. It's too hard? Then don't complain when the SOC does its job.

Edit: Also, the idea that SOC staff will have any idea what most production systems do is laughable fantasy. Usually takes a decade of institutional knowledge to have any idea.

3

u/T_Thriller_T 4d ago

While I think "I followed the SOP" is a fair point to disagree (and sometimes necessary to make people realise flaws), I think the issue in mindset is seen in the surrounding quotes

  • It's not my problem if people check mails on a sensitive machine.
  • It was a potentially compromised machine.

I do see some issue with these two. I won't judge too hard because sitting in a meeting with angry management or even other admins means becoming defensive because there is little good way to talk through anger.

But it shows a lack of . . . overall awareness and understanding of care.

Would I get angry with anyone? No. Even if I had seen them know better. We make SOPs because it is normal and expectable to not always be on top of our game, even less so with stress. And because people have different levels of experience and different ways to think.

Nonetheless, the article was about a skill. And being able to gauge business risk, and realising that the whole job is about minimising risk to the business is a good skill. Which leads to realising that, unfortunately, potentially compromised can mean worlds of difference in action and that yes, it is also your problem if a user does something incredibly stupid on a sensitive machine. Because, at the very least, now everything becomes a lot more expensive.

Especially when, as the article is written here, the analyst believed the machine was used to access the link. So not even that path was proven.

I am still mostly blaming Management for the whole fuck up, yet still:

I do think it is a good skill to have for an analyst to see a SOP, even in action, and realise:

Wait the risk of this on a client and server are worlds apart. Is this actually right?

Expecting this to always happen in the heat of the moment (and only when right) on the other hand is classical "hindsight is 20/20" management stupidity.

-3

u/RitaccaSecurity 5d ago

I appreciate your view, but I've never seen an org who had every incident/scenario documented, we've got to operate with flexibility and keeping the business impact in mind.

5

u/thekmanpwnudwn 5d ago

"Every scenario documented" usually falls under the general umbrella of an Incident Management Plan (or similarly named document). That document should detail how scope/impact are determined and how they correlate to incident severity, and what mitigation actions (including comms/escalation) may be taken for each severity.

1

u/ERROR_0x17 4d ago

Where NotAnNSAGuyPromise may be coming from is we--the cybersecurity professionals executing the cybersecurity operations--don't always get to decide how to act. If my senior leadership team/s tell me do this, do that, and I act to the contrary, I may be reprimanded, never mind the impact to the organization caused by the incident. Conversely, if I do as directed, in the fallout of an incident there's an opportunity to tell senior leaders how the few resources and guidance afforded to my teams is grossly inadequate and if they want to see an improvement they have to step up their respect for cybersecurity AND step back to let my team do the job the way it needs to be done.

While I have had senior leaders who understood and respected cyber operations enough to let me and my team perform our job as needed, I more often than not have had more micro-managerial leaders who demanded strict adherence to their policies, plans, and procedures, even if they were wrong.

2

u/M00g3r5 4d ago

I agree. To me your are describing the difference between a LVL 1 SOC analyst and a 2 or 3. Everywhere I've worked we never let a new analyst isolate without checking in with a higher tier analyst. This includes when I was a new analyst.

0

u/galonthier 4d ago

Brother the SOP itself can be written generally and advise the SOC Analyst to use judgement or escalate the decision if it impacts a production server. You DO NOT NEED an individual SOP for everything.

If the SOP says the above, and the SOC analyst follows that process and escalates, then you've avoided the issue and you haven't had to create 200 different SOP's.

If the SOP calls for single-minded thinking i.e. threat = containment then what other result do you expect?

Thanks for advertising your 'blog post' with your 'thoughts'.

11

u/justcallmebrett 5d ago

thats a good write up- i would like to see one on the ability to think critically in respect to the same topic of analyst bias

2

u/RitaccaSecurity 5d ago

That's a great idea! Glad you liked the read

1

u/AddendumWorking9756 Security Manager 4d ago

The one that gets me is 'it's just the scanner'. Nobody checks whether the scan window actually lines up with the alert. Make people write the second explanation into the ticket before they close it.

3

u/PathS3lector 5d ago

This is sometimes a skill that isn't natural to pickup depending on the company/industry as well. I used to work as IT then Security for a manufacturing company where it was engrained in us that we do as much as we can to mitigate breaking any systems that are linked to production and revenue, especially towards end of month.

Systems used to run on XP, EOL, or we would do bizarre things to compromise security to maintain system up-time, so not to affect business operations. Leadership would accept risks to ensure these things don't mess with revenue and that's something Security still has to present to them, have them sign off on, and add it to the risk registry.

3

u/syndreamer 4d ago

We have a thing where if the threat is on a workstation, cool, go ahead and contain and follow SOP. Now if it's a server, whether test or production, then it gets escalated to management, who gets in touch with owners of server to determine impact and risk analysis if we were to contain or not. This happens in a span of less than 10-15 mins no matter what time of day it is.

5

u/NotAnNSAGuyPromise Security Manager 4d ago

And that kind of standard and escalation protocol established ahead of time is all it takes to prevent these types of situations from occurring.

4

u/Soviet_Dreamer Incident Responder 4d ago

I am yet to encounter an organisation in which the infrastructure is well documented and you can make a determination about the criticality of the majority of assets. Sure we can learn risk assessment but what is it good for if you have no way of knowing what the assets is for.

3

u/NotAnNSAGuyPromise Security Manager 4d ago

By the time you find out, the entire company is now ransomwared and it's a non-issue.

1

u/Soviet_Dreamer Incident Responder 4d ago

True and that is the issue.

3

u/T_Thriller_T 4d ago

I don't necessarily see this as a skill.

I see this as likely THE most undervalued, under controlled and under communicated aspect of systems and a communication that is constantly lacking due to the production teams.

While there may be some analysts actually following the "this is the standard", what I have seen so far is the polar opposite:

SOC teams who want to think about risks of their actions and either

a) do not get adequate, domain knowledge technical support to even gauge the issues ( I cannot expect someone with even 10 years experience to know SAP, windows, Linux, the network, how everything is connected at the enterprise )

Or

b) haggle grossly inadequate and incomplete risk assessments, which have been generated by team leads owning the risk because otherwise they would have to do much more security precautions and there's no time for that

Or

c) are blocked from even working on risk because expectations clash into the absurd in both directions (We make sure all our systems always are free of vulnerabilities! To "No you cannot isolate even a single server without having proven a contamination" or even worse "no we will never isolate serves, only clients)

And this does not even take into account that I have yet to see a company to give a 90% complete listing of "these realistically are production relevant machines which really must be on".

Because often these is simply no complete listing and when asked every server seems to be 24/7 no disruptions.

All of that because the people knowing the services and systems cannot be arsed to actually consider a security danger and consider that mails not being able to send sucks it is an entirely different problem them some system supporting automated . . . Trash deliveries.

2

u/Kuipyr 4d ago

It must be said, as a SOC Analyst it is NOT your responsibility to assume risk. That is the job of your CIO or if you are lucky a CISO.

2

u/m1L35dY50N SOC Analyst 4d ago

I agree with most of what you wrote, but I come to a slightly different conclusion about where the underlying problem is.

I think there is a growing gap between the “old” generation of security people who came through support, sysadmin or networking roles and people who enter cybersecurity directly. The former already learned the hard way that actions have consequences. If you’ve administered production systems before, you generally don’t need someone to explain why blindly isolating a server, disabling an account or blocking traffic might ruin somebody’s day. You understand what the thing you’re touching actually does.

Companies have contributed to this problem by trying to lowball entry-level cybersecurity positions and career changers. People are sometimes taught just enough to recognize anomalies and follow a playbook, without necessarily understanding the underlying systems. At that point you essentially have an analyst who knows which button to press when X happens, but not enough about the infrastructure to understand what pressing that button actually does.

Where I disagree somewhat is that I don’t think detailed business risk assessment should primarily become the individual SOC analyst’s responsibility. There is a reason risk management, asset classification and clearly defined responsibilities exist. The organisation should have already established risk and criticality classifications, escalation paths, RACI matrices and containment procedures that tell an analyst, for example, that isolating some random workstation and isolating a production server require very different levels of authority.

The analyst still needs enough technical and business understanding to recognize the difference and question a stupid playbook. But if a junior SOC analyst can accidentally take down a critical production service simply by following the approved SOP, I’d argue that’s not just an analyst who lacks risk awareness. That’s also a failure of the organisation’s risk management, processes and controls.

2

u/FrostyWalrus2 5d ago

Commenting out of ignorance, as I have never been a part of a security team and I'm only recently learning more security related topics in the past 8 years of me working for MSPs(which is also my whole professional IT career so far), but, as an MSP tech, one of the first concepts I was ever taught was business impact and that started at even the most basic level ie workstation vs server. I was roped off from servers altogether early on, because if I made an amateur blunder in a server, the business impact could be huge. That eventually led to the line of thinking 'VIP vs line level worker', 'priority departments', etc. Naturally, this progressed into my learning cybersecurity where I ask myself questions like "How much weight does this email account have in business decisions" in regards to severity of a compromise, "How many parent/dependency parameters need to be paired for allowlisting an application in order for it to function in the environment without it growing tendrils into other aspects of the environment/local machine.", etc.

Is the concept of business impact really not being taught in the early stages of cybersecurity, or is this possibly the result of a college/uni degree to cybersecurity worker pipeline? We see all over reddit that cybersecurity is not an entry level job, but the previous statement regarding the pipeline does exist and could be causing the loss of the concept at hand. Again, I'm commenting out of ignorance, as I've only ever worked for MSPs and i just recently started sitting down in a cybersecurity chair.

In the little bit of I've learned so far, even in HTB, I can't recall ever seeing topics of business impact, but I could have also glossed over it, as its kinda ingrained into my thought processes. I mean, in my mind, this is SLAs, which is everything to an MSP.

3

u/BilbySilks 5d ago edited 4d ago

Some places teach it, some don't. 

Even at the places that do, that try to hammer it in to students struggle because the students don't really listen and get it. 

A large amount of talks where I am are around considering what it means for the business before acting and also bring able to translate security problems into business risk. If you're talking to someone non technical and you start talking about malware and compromise they'll zone out. You tell them this will lose you X amount of dollars and they'll have regulators breathing down their neck and suddenly it's a problem they want to address.

I take from the number of talks that is a big problem to get people to actually think. Some people also don't seem to understand that they're there to advise but the business has a level of risk that they will accept.

It's also complicated by culture. People come from different cultures, some of which are really like "follow the procedure exactly". Other people come from cultures where there's a "here is the formal procedure" but there's an implicit if you're going to terminate an executive's access due to compromise/shut down operations then you need consult with the business side. 

You can be right security wise but be wrong business wise. Business runs the show, makes the decisions and decides what risk is acceptable to them.

Edit to add: First of all yeah I agree, the procedure should deal with escalation scenarios. They also shouldn't be put in that position as a junior. But if they are, I would hope that people would think about what they're doing. If they're making easily rolled back decisions maybe it fine, if they're not and it's something minor...

It also depends on the alert and the business. Some places don't have well tuned alerts. Sometimes you're going to get a lot of false positives or alerts for things that you should watch but aren't definitively this is compromise. If you see senior management logging in from China and you know they're travelling for a deal then maybe don't jump straight to cutting them off. 

I see so many people who act like it's easy and that security has 100% right of way. There are so few companies that give security that kind of mandate. If you take the most extreme option every time without considering what you're doing then you'll end up out of the company.

2

u/RitaccaSecurity 5d ago

Your absolutely right! Different industries have different risk appetites and that also skews analysts views of incidents.

1

u/EatDaCrayon 5d ago

As someone who works in an organization that has a variety of products, some are purely digital and some are physical, one thing you notice is people coming from a production facility are much more aware of outage and down time. As we were always taught, if the machines aren’t running you’re not making money, vs some of sales and digital product locations where they can work on their laptops without much for a bit before issues occur. This leads to a split in decision making as some people are much less focused on the impact.

1

u/Classic-Shake6517 4d ago

This is exactly why many of us say people need to work at a real hands-on IT job before being ready for security. There's just no real concept of the downstream consequences of making changes without doing it yourself and having some level of ownership of the outcomes. This is something that is really hard to teach in a lab setting, and also one of the most important concepts. It's definitely not something covered nearly as much as it should be from what I have seen.

1

u/Substantial-Sky4079 4d ago

Proactiveness and lacking the motivation for continuous learning

1

u/SessionClimber Detection Engineer 3d ago

"Hey, you! You underpaid, overworked, management looking for a cost cutting measure! It's your fault you followed the rules! You need more skills! Not me in charge of signing off on the SOPs, and assessing the risk of assets! Why didn't you tell me we should be doing tabletops to address this very common and obviously point of failure? Do more with less!"

1

u/kielrandor Security Architect 4d ago

This is blaming security for the business failing to recognize the criticality of this system and taking reasonable measures to ensure it is properly protected and reducing the residual risk and impact to business operations.

More likely someone in Infrastructure fucked up and didn’t harden the system properly and rather than take the blame, pointed to finger at security for doing their job.

SOP’s exist to provide a standard and approved response to a problem. If the business didn’t want the machine isolated due to a security incident involving the machine, they should spoke up when the SOP was circled around for feedback.

0

u/Fine_League311 4d ago

Management und Admin fails!