CTIBench: Evaluating LLMs in Cyber Threat Intelligence with Nidhi Rastogi - #729

30 Apr 2025 · 56 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Notes on Episode #729 of The TWIML AI Podcast: CTIBench: Evaluating LLMs in Cyber Threat Intelligence with Nidhi Rastogi

Episode Overview

  • Host: Sam Charrington
  • Guest: Nidhi Rastogi, Assistant Professor at Rochester Institute of Technology
  • Topic: Cyber Threat Intelligence (CTI) and the development of CTIBench, a benchmarking framework for evaluating Large Language Models (LLMs) in real-world cybersecurity tasks.

Key Concepts Cyber Threat Intelligence (CTI)

  • Definition: Aggregation and analysis of cybersecurity-related information to detect and defend against cyber threats.
  • Importance: CTI combines various data types, including log data and threat intelligence reports, to anticipate potential threats in a network.

Evolution of AI in Cybersecurity

  1. Rule-Based Systems: Early methods focused on identifying patterns in network logs and files.
  2. Machine Learning: As data complexity increased, machine learning improved the speed and accuracy of pattern identification.
  3. Large Language Models (LLMs): Provided contextual understanding and the ability to generate informed responses based on extensive data.

CTIBench

  • Purpose: Evaluate LLMs' performance on cybersecurity tasks, focusing on their ability to respond accurately to CTI-related queries.
  • Benchmarking Task Types:
  • Knowledge-based questions (e.g., identifying attack patterns).
  • Reasoning questions (e.g., mapping vulnerabilities to Common Vulnerability Enumeration, CVE).
  • Practical tasks based on real-world situations faced by cybersecurity analysts.

Key Takeaways Advantages of LLMs in CTI

  • Speed: LLMs can analyze and correlate information much faster than human analysts (seconds vs. hours).
  • Contextual Understanding: They provide informed responses by synthesizing data from various sources.
  • Fine-Tuning: Techniques like Retrieval-Augmented Generation (RAG) ensure LLMs stay updated on the latest threats.

Challenges with LLMs

  • Knowledge Cutoff: Models trained on outdated data may miss recent cybersecurity information, leading to inaccurate responses.
  • Hallucinations: LLMs may produce convincing but incorrect information, particularly in high-stakes domains like cybersecurity.
  • Complex Task Performance: Smaller models tend to struggle with complex tasks, confirming that model size impacts performance.

Benchmarking Process

  • Data Sources: Utilized trusted standards like NIST and MITRE for building the benchmark.
  • Question Generation: Employed ChatGPT to create questions, followed by human evaluation to ensure accuracy and clarity.
  • Evaluation Method: Responses were manually verified against real-world cybersecurity data to assess LLM performance.

Future Directions

  • Mitigation Techniques: Research on how LLMs can assist with threat mitigation.
  • Explainability in AI: Exploring how LLMs can provide rationale behind their conclusions to enhance human analyst decision-making.
  • Concept Drift Monitoring: Addressing how models can detect changes in threat patterns over time and when they require retraining.

Conclusion Nidhi Rastogi and her lab’s work on CTIBench marks a significant step towards understanding how LLMs can be applied effectively in cybersecurity, highlighting the importance of benchmarks in identifying model limitations while also exploring future opportunities in AI-driven threat intelligence.

Additional Resources

  • [Complete show notes for Episode #729](https://twimlai.com/go/729)
  • Future research insights from Nidhi Rastogi's lab on LLMs in cybersecurity.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We've identified a space where there is no such existing benchmarking model which can tell whether this specific language model is capable of giving good responses or accurate responses on cyber security. specific tasks. And if it is able to, is there a metric to determine how well it is performing on those tasks, you know? So we designed something based on what an actual threat analyst would experience in a given day.

0:44All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Nidhi Rastogi. Nidhi is an assistant professor at Rochester Institute of Technology. Be sure we get going. Be sure to hit that subscribe button wherever you're listening to today's show. Nidhi, welcome to the podcast. Thank you, Sam. Thank you for having me. I'm looking forward to our conversation. This one has been kind of a long time in the making. I think we first connected many, many years ago, and I think it's a good time for us to talk because the space that you work in, cybersecurity and security in general, and of course, AI have both been moving at an incredible pace.

1:27We'll be talking about the intersection of those two areas. But I'd love to have you share a little bit about your background. Sure. So, as you mentioned in my introduction, I'm an assistant professor at Rochester Institute of Technology, and I've been at RIT for the past four years, working at the intersection of AI and cybersecurity. But the genesis of me getting interested in this space was long ago when I started my PhD at RPI. And because... Go engineers. Oh, yes. and air was just picking up it was in the name of machine learning back then however before I did my PhD I was in industry working in Verizon Wireless and GE I had some internships at Yahoo and IBM so these kind of you know created this mindset for me how to approach research which is not just purely theoretical research but also thinking of it from a practical perspective.

2:30So I decided to do a PhD in more like applied research in cybersecurity and along the way I picked up machine learning. And that's what I'm doing right now. I'm very, very excited about it. I've always been passionate about cybersecurity. And AI has kind of transformed how threats are detected, how we defend organizations, large networks against threats using AI. So that's what I'm busy doing these days. I'd love to have us start off by talking about how LLMs have changed this intersection between cybersecurity and AI. I remember, we've talked about cybersecurity a few times on the podcast, several times.

3:17And the last of those times, I think, was kind of in the deep learning era. And we talk about how, you know, networks are trained to kind of identify patterns in, you know, either logs or in software in the case of like antivirus. How have LLMs changed the game? Yeah. So let's talk about it from the perspective of a before and after. After means what LLMs are doing and before is what purely machine learning and deep learning is capable of doing. So even before that, it was just rule-based determination, you know, finding patterns in network logs and in some kind of files behaving in some anomalous manner and actually malicious manner, just identifying what pattern it is.

4:09But the size of the data used to be much smaller back then. And, you know, as data kept growing and the complexity of files and, you know, the network logs and everything started to increase, machine learning became even more important because we needed to identify those patterns, not just accurately, but also fast. And that's what machine learning and deep learning were capable of doing. And exactly where my PhD was focusing on earlier. However, the challenge became slightly more when understanding, comprehending what these models are deciding or decisions they are making. They required getting trained on not just simple network logs and files behaving maliciously or in a benign manner.

4:59They also needed to be trained on other kind of modalities of information, like threat intelligence reports, so that the model is not only able to give you a decision, but also context behind that decision. And that's where large language models began to play a more prominent role. And today what the output, which is the after, is what the output we are able to see is we are ingesting small information, you know, some kind of a query into these LLMs as a prompt. And these LLMs are able to gather from the context and, you know, associate that with the response. And now what we get is a lot more informed response from AI or LLMs, so to speak.

5:46So this is the kind of the journey that we have seen is purely giving the decision. The model is just purely making the decision. But now there is context behind it. And we can continuously query the model for more information and, you know, make a decision based on that. And is the basic pattern that the log information is put into the context of the LLM? And I guess I'm asking this from a couple of perspectives. One is the increase in context windows over the past year, let's say, must be a huge boon to any kind of log processing and cybersecurity, as well as I'm also curious about the application of retrieval oriented types of paradigms.

6:37Like, is there a rag kind of analogy in cybersecurity? Absolutely. So if you look at regular large language models, you know, like a chat GPT, no, I wouldn't take that as an example because the performance of chat GPT is much better. Let's say an open source model like LAMA, 7 billion parameters or 70 billion parameters. It has a lot of information that was used to train it. but if I get a risk but the problem with these trained models is that there is always a cutoff date when it was trained on so let's say a 70 billion parameter lama model was trained on 2024 data set and if I query it today which is you know like in the second quarter of 2025 if I'm querying it at that time, it may miss out on the recent cybersecurity information.

7:37Let's say there was a malware which was recently discovered like a couple of weeks ago. That information would not exist in a model which had a cutoff date in 2024. Now, at that point, how do we make sure that this model has more recent information? That's the place where rags come in. We can fine-tune these models using, you know, some more recent information using a RAG approach. And then these models can be up to date with the more recent information. So that's exactly where RAGs come into play. Your work is focused on an area of cybersecurity called cyber threat intelligence. Can you break down what specifically that aspect is focused on?

8:21Yeah, so cyber threat intelligence is basically an aggregation of all of cyber security related information available on the Internet, aggregating it, analyzing it, and then using it to detect any kind of cyber threat as well as defend from those cyber threats. And why that is important is because, like I mentioned in the beginning, cybersecurity information is available in all kinds of modalities. There isn't just log data, you know, where there is some malicious threat patterns detected or there is some virus which has been sitting around in a file. There are also these reports which are written by security experts sometimes, sometimes by journalists, sometimes by people just purely interested in cybersecurity.

9:16They write these reports about what they have recently found in the cyberspace or in some specific organization. Let's say there is an enterprise which has recently seen an APT attack. They will be able to describe the detection, even the detection of the attack, what are the different steps that were followed called tactic techniques and procedures, which were followed by the attacker and what they were able to identify. All of that is described. They give even specific information like indicators of compromise or the hashes of the malware. All of that is contained in these reports. These reports can go, you know, up to like 50, 60 pages long, or they could be present in the form of a very small tweet or blog reports.

10:09So cyber threat intelligence is basically gathering all of this information along with telemetries, which are telling you the, you know, the log information in a network. All of that comes together. That is called cyber threat intelligence. And what we are doing is we basically gather all of this information and kind of make predictions. Given all of this data, is it telling me anything about potential threats present in my network? Or if there is an existing threat in my network, what are the next steps that this threat might take? So these are called attack patterns. What is the next pattern I might find in my system, in my environment?

10:56That's a space that we look into in cyber threat intelligence. Can you talk about some of the ways that LLMs are used in CTI? Yes. So usually what the cyber threat intelligence analyst would do is they would gather all of this information and then try to find some kind of existing pattern in their network and then associate that pattern of anomalous activity or malicious activity and relate that to the cyber threat intelligence report. Or they may, you know, vice versa. They find something in the threat intelligence report and then might see, oh, this is relevant to my environment. So let me see and find something existing in my environment.

11:40So they kind of correlate what's inside the network and what's out there. The advantage of LLMs is that all of this information gathering and analyzing and predicting can be done on the fly. You know, they are trained on a lot of world knowledge. They are trained on specific cybersecurity domain knowledge. So they already know about a lot of the threats that are out there. And then if you have implemented a RAC, then they're fine-tuned with that recent information also. So what an analyst would have taken a couple of hours to assess, to analyze and to comprehend and then associate with a threat pattern they have identified in their network, LLMs are able to do that in a couple of seconds.

12:29So they can summarize all of this information and at the same time also connect, well, how is this information relevant to your specific organization? but that doesn't mean that there are no challenges as you know we've heard it one too many times if an LLM is not up to date then it will not give you the correct threat report or a threat information will be missing if it doesn't know about what's going on what's going around at that time and then there is obviously this challenge of hallucinations if it doesn't know then it might just give you a response which might not even apply. And in domains like cybersecurity, this can be, you know, kind of detrimental.

13:20I may be getting a response which very, very, you know, it may sound very, very convincing. And if the analyst doesn't know, they might fall for it and, you know, end up taking measures or taking steps which may not benefit the organization. Maybe taking a step back, Can you talk a little bit about the kind of broad applicability of LLMs in a sense of LLMs have been trained on, you know, primarily written, written word and, you know, log files, while they can have written words in them, they're like not necessarily following the patterns of language. And so these things that are trained to generate, you know, pros essentially and identify patterns in pros, you know, might not necessarily be the best thing for, it's not a foregone conclusion that they're going to work great in log files.

14:21Can you talk about kind of what we know about how LLMs perform with log style data or cybersecurity style data? That's an excellent question, Sam. So 2023 December is when we first figured out that LLMs are able to comprehend log data because LLMs were just coming about. And I had a master's student visiting us from Germany, and she was interested in LLMs and knowledge graphs and cyber threat intelligence. And we were like, let's just give it a shot and see what an LLM comes up with. And we were very, very surprised. Not only were they able to comprehend that format, But they were also able to identify that there is a presence of a malicious activity in these logs.

15:07And this was our first take on just exploring LLMs. And we were very surprised by that. I also need to mention here is the capability of the language model is very much determined by the size of the corpus, the data corpus that was used to train on it. Initially, it was LAMA 7. Like the 7 billion parameter model? Yes, exactly right. So that was the one which was available to us back then. We were very surprised by the quality of the output. I wouldn't say the quality was high, but it was able to comprehend. And that's all what we cared about at that time. So yes, language models are able to comprehend a vast variety of formats and syntax that we wouldn't know.

15:57it has been trained on, but just give it a try and it'll surprise you. We've used different kind of languages, coding languages, scripts, JSON, log files, even hashes of malwares. And it was able to associate that. And I'm talking the open source Lama 8, 7.8, which has 7 billion parameters, 8 billion parameters, and even Lama 7.8. We have trained it, we've explored all of these, and they've surprised us with how many variety of syntax it is able to comprehend. And I guess once you start asking these questions, that leads you naturally to, you know, how well are these LLMs doing, and how do we know how well they're doing, and what's a benchmark for LLM performance in this space?

16:48Yes, so our work is specifically only being in cyber threat intelligence. And that is what we have benchmarked different language models on. So we've benchmarked LAMA 7, LAMA 8, the 70 building parameter, LAMA, ChatGPT 4, ChatGPT 3.5, Gemini, and many, many others. But these are the more popular ones. And as I mentioned earlier, the more robust and large the corpus, the training corpus is, it is kind of a very big factor in the quality of the output that one gets from these language models. So 7 wouldn't perform very well on slightly complex tasks. It would perform very well on simple, straightforward tasks.

17:39But a chat GPT-4 is what we were using back then. Now we have 4.5 and many other versions. But 4 was performing exceedingly well, especially in the cyber threat intelligent tasks. And there were many tasks that we used to benchmark all of these models. However, there were times when they were not performing well. So there were specific tasks they were performing very well on, almost three to four, three out of four times. The more advanced models, their performance was very, very, very, very high. But the smaller models were doing only, you know, so, so much, able to do only so much, like 50 % of the times on even simpler tasks.

18:26And so we've started talking about one of your research projects or one of your recent projects called CTI Bench, which is a benchmark for evaluating LLMs in cyber threat intelligence. And you mentioned a number of the LLMs that you benchmarked against this task and data set that you created. Are there also like specially trained cyber threat intelligence LLMs that people have created beyond the Llamas and the ChatGPTs and Geminis? Yeah, so the ones that I've talked about are general purpose language models. For the cyber threat intelligence space or cybersecurity overall, there is one called Sec Gemini version one, which came out of Google.

19:27That is the one that we are aware of. There may be others, but for cybersecurity specific, this is the one. And they also used CTI Bench to benchmark their security model. Can you talk through some examples of CTI Bench? Like what are the tasks associated with it? Sure. So CTI Bench is basically a benchmarking framework. You know, what it does, we kind of identified a space where there is no such existing benchmarking model, which can tell whether this specific language model is capable of giving good responses or accurate responses on cybersecurity-specific tasks. And if it is able to, is there a metric to determine how well it is performing on those tasks?

20:21You know, like just the way we are graded in school or in college, you know, you've gotten like eight out of 10, for example, on this kind of, you know, MCQs or on subjective questions or on coding questions and so on and so forth. So we designed something based on what an actual threat analyst would experience in a given day. So a threat analyst like a junior or a slightly experienced analyst would experience in a given day. So it was very, very rooted in practical applications. So having said that, the questions that we, the tasks that we assigned for CTI Bench were, we kind of organized, again, based on the actions or the activities that the security analyst would be pursuing in a given day.

21:14So these could be knowledge and reasoning in cybersecurity. So knowledge would mean how much do you know, like textbook knowledge? Do you know about a pattern? Do you know how to find the ID of this pattern? Which sources should you be looking up if you need more information? For example, I'll give you a very simple example. What is MITRE attack patterns? If a cybersecurity analyst is asked this question, they would tell you it is, It basically is a threat attribution. You know, it basically, it's a website where MITRE has aggregated all the threat attributions and tactics and techniques and procedures and so on.

21:59So we were kind of querying, we kind of query these kind of knowledge-based questions on one type of tasks. And then there are reasoning questions like if there is a description of a vulnerability, of a software vulnerability, will the LLM be able to, will it be able to know which pattern CVE, which is common vulnerability enumeration, which kind of is an ID to every kind of software vulnerability out there. Will it be able to identify the correct ID of that? And we kind of designed those kinds of questions, knowledge-based reasoning questions in the form of MCQs, in the form of threat attribution, in the form of attack pattern determination.

22:50So we kind of classified four to five types of tasks in CTI Bench. And each of those tasks had questions. the LLM was to respond to those questions and then we would evaluate the LLM on those questions to determine whether it is performing well or not. And when I think about the types of questions that you described in those tasks, they're primarily, I guess I would think of them as meta questions, They're kind of facts and tidbits that knowledgeable cyber intelligence analysts might know, but they're less related to, you know, the mitigation or detection of one of these threats. Does the benchmark cover that at all?

23:38So that's a very good point. Initially, the tasks that we were focused on for this paper were, for example, root cause mapping. You know, there is a software vulnerability, like I said earlier. Is the LLM able to map it to an existing vulnerability and tell us the ID? So there's a root cause mapping. And then the next one is like threat attribution. Is the LLM able to, based on a snippet of threat intelligence information, which is a text, if I provide the LLM a snippet of the threat intelligence, is it able to tell me what malware family does this threat belong to or who is attributed to or which threat actor is responsible for this kind of tactic and technique?

24:28So these were the questions. So you are right about it. We are more in CTI bench, curious about attribution, curious about knowledge, curious about how well is the LLM able to understand information and able to produce an output which is helpful and how accurate it is. But when it comes to mitigation or remediation, where instead of just providing there is a vulnerability, we are also providing another piece of information, which is how to mitigate this vulnerability. And because this kind of a pattern has been seen in the wild, because it belongs to this threat actor. So essentially connecting the dots and at the same time, also telling you what are your remediation measures.

25:16We do not do that in CQI bench. However, this is the focus of an ongoing research in my lab. So in some ways you can think of it as like CTI bench is analogous to when we benchmark the NLM against like the bar exam or the, you know, the MCAT exam. It's like the background knowledge that someone working in this space might have, but not necessarily how to go in and solve particular problems. That's work that's ongoing. Yes, that is correct. That is a very accurate analogy. Can you talk a little bit about the process of building out the benchmark? Yes. So CTI Bench, like most domain-specific language models, it requires a trustworthy source of data.

26:09And then once you have the data and you have kind of told the model about it, you want to generate some set of questions around that data. because you want to get very specific responses, accurate responses, especially in a domain like cybersecurity. Accuracy is very, very important when you're measuring the performance of a model. And then once you have created those questions, we want to get responses on real world data. So it's like, I have the knowledge, a textbook knowledge. I have my question paper. Now I want to test you, but on real world knowledge. So I've kind of trained my model, but I want to test it on real world knowledge.

26:55And this is all, again, at the core of our research, which is very applied, very rooted in practical application. So we don't want to create a toy language model or something which is hypothetical, which may not be used by anybody. And therefore, these were the three foundational steps when we designed CTI Bench. So for the source material, for example, we relied on NIST standards. We relied on MITRE attack patterns. We relied on GDPR because it has security and privacy regulation related information. So these were the source of information that we basically pulled from. And these are also familiar foundational work for us.

27:48And then when we were designing the question, and this was the most interesting part, we used ChatGPT4 to create questions based on the sources that we pulled. We basically asked, can you create questions, like MCQ questions with these kinds of responses, four responses, for example. And it did generate 500 and then 1 ,000 and then 1 ,500 and so on and so forth. So we kind of generated 2 ,500 questions from ChatGPT4. However, it was not, you know, just like most LLMs, we could not completely rely on it. It sounded very convincing, the questions and responses. But that's where a lot of the effort went in, is we kind of reviewed every single question and the responses.

28:36And we had to make sure that the responses are not confusing or we shouldn't have more than one response. or the questions should not be framed in a way that it can confuse the reader and, you know, they may not be able to pick a correct response. So all of that effort went into designing, you know, right set of questions and ensuring that the responses, there is one proper response for every MCQ. And then for the real world data, which we used to evaluate the questions or, you know, all of these different language models, We relied on cybersecurity-specific information like CVE, which is the vulnerability enumeration.

29:20It describes every single vulnerability out there. And also the weaknesses in a given software or in a given environment. All of these are available on MITRE websites. So these were the three sources of information and how we created the data. Now, once all of that has been done, we kind of designed tasks around all of this information. So the tasks, like I mentioned earlier, are knowledge-oriented questions. So knowledge-oriented questions were the MCQ questions, the CTI MCQ. Then there were practical tasks like vulnerability mapping or threat attribution. And then beyond that, we also wanted to see if the LLM is able to compute severity scores from text information.

30:07And that is called CVSS, calculating vulnerability severity score. So all of these tasks are important to a threat analyst. MCQ is basically knowledge retrieval, recall, remembering important security information, vulnerability mapping and threat attribution. All of these are important tasks on a daily basis, which an analyst is required to perform on a daily basis. And then CVSS is some kind of a severity score calculation that the NEMAs need to do to be able to triage a given threat. If it is important, very important, or critical, all of that is based off of the CVSS value. The higher the severity, the higher the score.

30:55Is CVSS an existing industry kind of formula? Yes. Or, okay. All of this are based on industry standards, what most of the organizations around the world use. So none of this was invented by us. Got it. Got it. Got it. can you talk a little bit about the process of then rolling this out and kind of testing existing LLMs any surprises there a lot of surprises as was expected these are language models but like any model there is always room for surprises so how we roll this out was after we trained the models, sorry, find you on the models. And then once we identified the tasks and we had these responses and everything, we started to subject all of these language models to these questions and to these tasks.

31:55And what we learned, and very unsurprisingly, is that ChatGPT4 performed exceedingly well on questions such as MCQs. And ChatGPT4 was the latest and greatest back then, which is mid of 2024. That is the timeframe when we submitted this paper at NeurIPS. And then a very close second was an open source model, which was LAMA 70 billion parameters. It was able to perform very well, not as well as chat GPT-4, but on most of the questions like MCQs and thread attribution and attack pattern identification and CVS scoring, it was able to perform very well. So that kind of gives us kind of a hope, you know, that we do not need to rely completely on expensive closed source models such as JGBT4 or their latest versions.

32:51So the smaller models, LAMA7 and LAMA8, again, that is the billion parameters that was used to train them, they performed only well on simpler tasks. And very interestingly, pretty much every single model fared very poorly on some of the questions. And that was very interesting for us. Why? Because we kind of looked back at those questions. What is it about the questions? Is it about the training? And then we also had human evaluators who were also experts kind of checking if the question is complex or does it require a lot of detailed understanding or deep knowledge, which is what we determined was the case.

Read the full transcript

33:39And are you saying it was the same questions tripped up all of the LLMs or each LLM had their own set of questions that it struggled with? Yeah, so some questions every single model tripped. You know, they weren't able to respond. And there could be many reasons behind it. I wouldn't say hallucination is the only reason. There is also the cutoff date. When was the model training cutoff? And that is a testing that we did was we did not expose the model to the last two months of the data or last few months of the data. But we tested them on those data sets. And either it gave a very, you know, informed response, well-educated response, still inaccurate, or the model was hallucinating, or it came kind of gave something that it was not trained on.

34:33So there were different reasons why the models were not giving correct responses. And we did dig a little deeper into that. And as someone who's building a benchmark, are those types of questions, are those your friend or are those, you know, types of questions that you don't want to see when you're trying to produce a benchmark? It's a very interesting question. And I, and this is something that we also mentioned in the paper is that benchmarks are kind of an approach to telling what the model is capable of doing and where it will fail. and identify those, you know, those blind spots or those edge cases where it either needs better training data set or it is, or there's something else needs to be done.

35:23You know, turning a blind eye towards it is not going to help anybody. So it's better to be aware of that. And that's what the role of the benchmark is. It's not to hide any kind of these edge cases or corner cases, but to reveal them so the analyst knows that my model will be able to respond 80 % of the time for these type of questions. And for the other type of questions, we need more of a human in the loop or some kind of human intervention or an expert intervention so we're not making mistakes that might cost us. Or what you see often with benchmarks is, you know, in year one, models perform, you know, 20%, 50%, whatever the percentage is on these questions, but then year two and year three, and as the models mature, they're able to solve more of these challenging, you know, these problems that were previously challenging to them.

36:19Have you continued to test newer models since the, since CTI bench was published? We haven't, you know, continued down that path yet. But what we are interested in is more of a mitigation effort. You know, how do we resolve or how do we find solutions to mitigating a threat once it has been identified? When you're formulating these tasks, you mentioned a good portion of them were these MCQs, multiple choice tasks. Were they all multiple choice or were there a portion that you needed to evaluate in more complex ways? There was knowledge oriented questions, which was MCQs. And then there were questions which were, you know, the CVE mapping to the CWE.

37:13So what we did was provided a snippet of the threat. And then now the expectation from like two of the tasks was the snippet that we provided, we kind of very carefully hid any kind of clue where the model is able to conclusively determine that this snippet is coming from this CVE or this CWE or which threat actor is responsible or behind us, this kind of a threat. So we very carefully removed that information from the snippet that we provided. And the outcome of that was basically the model mapping to the CWE and giving us the identifier, number one. And in the case of threat attribution, it was able to tell this threat could be the APT28, for example.

38:10could be responsible for this threat. So these kind of these kind of responses we evaluated, sorry, we verified using human analysts and that is what we did for all the MCQs. The questions, the question, the choices as well as the responses. We verified everything. But in all cases, there was one and only one correct answer and it wasn't like you were We're trying to evaluate a text description of a vulnerability and needed to use like an LLM as a judge to determine whether it was correct enough or anything like that. Yes. So aside from MCQs and, you know, identifiers and CBS score, the other thing we did was attack pattern identification.

39:03That was one of the tasks. So attack patterns are basically, these are like IDs. Again, there is an ID and then there's a description. So when an attack takes place, it always takes place in steps. There's step one, step two, for example. Initially, there's a reconnaissance. And then there is some kind of privilege escalation. And then there is some command and control server accessing your system remotely. And then there is a possibility of information extraction or data exfiltration. So these are the kind of steps, for example, that we see when an attack takes place, especially through an APT.

39:48What might happen is that we have in our environment, for example, we have seen initial stages. We have seen some kind of reconnaissance happening. There is some evidence of a bad actor present in our network just by the way of some logs have revealed that kind of a pattern. So all of that gets documented in the threat report, but it's kind of very, very sequential. When we have that kind of a threat report, we pick an initial part of the threat report and provide it to the LLM. and we ask the LLM, what are the rest of the threat patterns? And it provides us with those threat patterns along with the ID.

40:31Now, those IDs are very, very standard and are part of the MITRE ATT &CK patterns framework. And then the LLM is able to provide those IDs as well as a description. Because we had those threat reports, we kind of gathered 50 threat reports from MITRE website, and we only use those standard reports to determine if the LLM is able to tell us what are the steps that this, you know, threat vector could follow. We only provided the first few and then it was able to tell what the rest of the patterns would be. And then how did you find the process of, or how did you then evaluate whether it got those steps correctly?

41:16That was done manually. Oh, it was done manually. Yes, those were done manually. Yeah, all the questions, we had about 50 questions. We had a small team, but very active, hardworking team. We basically verified every single response. And we also provided them with the threat report so they could, you know, verify what is in the threat report is the same as what the LLM is responding. Did you come across any interesting examples of LLMs hallucinating in trying to run the benchmark? Yes. So that happened depending on which LLM we are referring to. So sometimes the response would be, you know, the attribution would be incorrect.

42:01So in the sense that instead of APT28, the LLM is saying it is APT12. So what APT28 and what APT12s are, these are standard ways of referencing to some threat actor. Like APT28 refers to fancy beer. So instead of saying fancy beer, we'll just use APT28. So it's like an identifier. So sometimes it would, you know, sometimes if you would ask a question, like I said earlier, is what is the purpose of MITRE attacks? So instead of saying that it is like a repository of adversarial tactics, techniques and procedures, it kind of catalogs all of them. Instead of saying that, the LLM would respond to, it is basically scanning vulnerabilities.

42:51That is the purpose of MITRE attack, which is actually true because MITRE as a company does that also, but not MITRE attack, which is, you know, another aspect of what MITRE does. So that would be, you know, kind of misunderstanding or be termed as hallucination. but the real hallucinations would happen when it was misattributing the threat instead of fancy beer it would attribute it to somebody else like aptx and at the same time it would say it very convincingly adding information you know like this threat was found in this organization at this location on this in this time period in this country and so on so it would be very very convincing but incorrect.

43:38But sometimes it would just make up information, you know, like saying cross-site scripting is related to SQL injection, but in reality, for that specific example, it was just improper input validation. So some kind of mix-up we would see, but very, very convincingly. So we saw all of those examples when even the best of the best LLMs were especially Charged GPT, they would start to hallucinate. The benchmark, you mentioned the importance of the knowledge cutoff for the benchmark. Does that mean that the benchmark and the benchmark incorporates kind of facts up to a certain day? So like the LLMs have a knowledge cutoff, but the benchmark kind of has a knowledge cutoff also.

44:32Are you going to continue to update the benchmark to incorporate the latest CVEs and vulnerability notes? So essentially what a benchmark does is we provide like a framework of how to evaluate a model for the specific domain. And if the model is up to date, we can add more questions to the benchmark. but that wouldn't require creating a new benchmark. That would be, you know, one can personalize that, let's say in two years or in three years. However, different kind of threat vectors will come about. What we are talking is only cyber threat intelligence. There would be potential attack vectors coming from AI that we did not include in this benchmarking.

45:24So yes, that is a big miss. And then the other miss is how do we mitigate? There is no benchmarking of is the mitigation techniques that was provided, how accurate are those? So some of those aspects we definitely miss out on, which is what I think gives ideas for future research. Of those two, it sounds like the mitigation is your focus for going forward or you kind of have research bets placed in both of those areas? Certainly for mitigation. but then there is also another aspect of this entire cyber threat intelligence and LLMs is aside from benchmarking, are there other ways of getting into the head of the LLM?

46:10Can the LLM explain itself? Can the LLM explain why did it give this response? That takes us to another side of my research, which is explainable AI. Can the machine learning model or AI model explain itself along with the confidence that it has in its response? So that is another avenue of research that we are pursuing for not just LLMs, but also for pure AI, like machine learning models trained on a certain corpus. Can the LLM or the machine learning model be confident about the response and how much confident it is? so we can determine when human in the loop should take place. And then the mitigation effort is also very, very important in this aspect because MITRE, again, provides mitigation techniques along with the attack tactics, techniques, and procedures.

47:02So we do have ground truth available for that. So mitigation and how SOC analysts can optimize their time when LLMs are, you know, treated more like an assistant. So how does that new environment look like? We are definitely looking into that space. But what was the most surprising thing you learned in the process of building out the benchmark? So it was definitely the need was, you know, something which was very surprising to us that there is so many benchmarks getting built or designed every other day, but there's nothing for cybersecurity, which can be called applied and practical. The ones which were out there were not very useful in practical scenarios.

47:51So that was obviously a gap that we identified, that my research team identified. But what was kind of not surprising, I will say, but we validated is that even the most well-trained models will have, they will have blind spots and we need to identify them. Because, again, we need to integrate these systems into our environments. For that, knowing exactly where it will perform, you know, poorly or where it will fail, especially in a mission-critical, you know, use case like cybersecurity, it is very, very important. So it was good to know that even the best of the best models are able to fail sometimes.

48:36And also what was very interesting was that the large language models, which are open source using high parameters, they're not a complete disappointment. They're actually very, very useful. So those are the two things that kind of was very helpful conclusions for us. And any particular challenges that stand out in the process? The challenge I will say was, I think just training all of this was very, very cost intensive. We did end up spending a lot of, you know, training time and evaluation time and everything. It did cost us a lot of money. What specifically was trained? No, I mean to say requesting tokens from GPT.

49:28Oh, like inference. Yes. Got it, got it. And then we also explored rags, which we did not include in the paper. That was extremely, you know, demanding in terms of the number of token requests and the response. All of that was, you know, kind of prohibitive for, especially for a research, you know, research lab. We do not have a lot of resources in our hand in terms of how much can we spend on these LLMs. So that was the biggest challenge and probably also the reason why we haven't come out with many such benchmarks because it did cost us a lot of money. You mentioned that Google published, is it SecGemini?

50:15SecGemini version one. Yeah. What was your observation about that model? So SecGemini was not available at the time when we published this work. And it came about, I think, around the same time, just a couple of months later on. They did use our benchmark to evaluate the security specific model, the LL that they came out with, with other general purpose models. and on the specific tasks, such as the CTI MCQ, which is a multiple choice questions and then threat attribution and also root cause mapping, which is basically if you provide a description, is it able to map it to a CVE? And then that CVE can be mapped to a specific weakness in the software.

51:04So especially on these two tasks, these are the tasks that were mentioned on their blog. I don't know about the rest of them, but apparently SecGemini version one performed quite well compared to every other general language model. Which kind of also tells us that if you want your language model to perform well on a specific specialized task, then they need to be trained on that kind of data or need to be fine-tuned on that kind of data. So that's what I observed from the results. And then we talked a little bit about kind of future directions based on this research. But can you talk a little bit about what your lab is doing broadly?

51:54Is it primarily focused on the CTI work or are there other aspects of security that you're looking at? So there are some other aspects of cybersecurity that we are looking into. But I'll start with cyber threat intelligence. So cyber threat intelligence using LLMs, but also mitigation techniques, how the LLMs are able to produce reliable mitigation techniques. We are also looking at concept drift in detecting threats, which basically means that if a model has been trained up until a certain point and there's a cutoff date, will that model be able to detect or be able to identify some kind of anomalous malicious behavior in malwares or in you know in the network through intrusions because training and retraining a model is not always possible it is a very very time and compute intensive process so what are the best ways of identifying when has the model begun to drift and that could be a language model that could be machine learning model.

53:01Just identifying when does the model begin to drift is crucial in determining when does it need to be retrained or when should it be fine-tuned. And then we are also working on proposing methodologies for retraining the model or fine-tuning the model so that the drifted model has caught up to the recent threats. So that is one aspect of the research we are doing. I also mentioned about explainability, which is, you know, kind of a need of the hour because AI is everywhere. But for specific use cases, we cannot be comfortable just with the response or the decision that the AI is making. We also need to know the rationale or the reasoning behind the decision.

53:48So we are working extensively in this space. And the use and the applications where we are applying explainability is in autonomous vehicle security, as well as in the SOC environment, which is security operation centers, where the analyst, the end user, the SOC analyst, is able to not just, you know, work with the alerts, because that's the job that they do, is working with true positives and false positives, but also learn about why is this alert called, is it a false positive? Or why is this alert, a critical alert? So kind of giving evidence along with the alerts. That's another area of research that we are doing.

54:36Very cool. Well, Nidhi, thanks so much for jumping on and sharing a bit about what you and your lab have been up to. I'm excited. this was really very important work for us because we do so much extensive work in this space and then the students were interested in exploring LLMs how much an LLM is able to perform well on cyber security tasks like CTI tasks just through pure grit and determination we were able to pull off this NeurIPS benchmark paper within a week so I would like to give a shout out to my students with that. That's amazing. Awesome. Well, thanks so much, Nidhi. Thank you so much, Sam.

55:19Thanks for having me.

55:43Thank you.

From the publisher

Today, we're joined by Nidhi Rastogi, assistant professor at Rochester Institute of Technology to discuss Cyber Threat Intelligence (CTI), focusing on her recent project CTIBench—a benchmark for evaluating LLMs on real-world CTI tasks. Nidhi explains the evolution of AI in cybersecurity, from rule-based systems to LLMs that accelerate analysis by providing critical context for threat detection and defense. We dig into the advantages and challenges of using LLMs in CTI, how techniques like Retrieval-Augmented Generation (RAG) are essential for keeping LLMs up-to-date with emerging threats, and how CTIBench measures LLMs’ ability to perform a set of real-world tasks of the cybersecurity analyst. We unpack the process of building the benchmark, the tasks it covers, and key findings from benchmarking various LLMs. Finally, Nidhi shares the importance of benchmarks in exposing model limitations and blind spots, the challenges of large-scale benchmarking, and the future directions of her AI4Sec Research Lab, including developing reliable mitigation techniques, monitoring "concept drift" in threat detection models, improving explainability in cybersecurity, and more.

The complete show notes for this episode can be found at https://twimlai.com/go/729.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
CTIBench: Evaluating LLMs in Cyber Threat Intelligence with Nidhi Rastogi - #729The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 56 min
Listen in VO