Considerations for Evaluating Large Language Models for Cybersecurity Tasks
• SEI Report
Publisher
Software Engineering Institute
Abstract
Generative artificial intelligence (AI) and large language models (LLMs) have taken the world by storm. The ability of LLMs to perform tasks seemingly on par with humans has led to rapid adoption in a variety of different domains, including cybersecurity. However, caution is needed when using LLMs in a cybersecurity context due to the impactful consequences and detailed particularities. Current approaches to LLM evaluation tend to focus on factual knowledge as opposed to applied, practical tasks. But cybersecurity tasks often require more than just factual recall to complete. Human performance on cybersecurity tasks is often assessed in part on their ability to apply concepts to realistic situations and adapt to changing circumstances. This paper contends the same approach is necessary to accurately evaluate the capabilities and risks of using LLMs for cybersecurity tasks. To enable the creation of better evaluations, we identify key criteria to consider when designing LLM cybersecurity assessments. These criteria are further refined into a set of recommendations for how to assess LLM performance on cybersecurity tasks. The recommendations include properly scoping tasks, designing tasks based on real-world cybersecurity phenomena, minimizing spurious results, and ensuring results are not misinterpreted.
Cite This SEI Report
Gennari, J., Lau, S., Perl, S., Parish, J., & Sastry, G. (2024, February 20). Considerations for Evaluating Large Language Models for Cybersecurity Tasks. Retrieved August 19, 2026, from https://www.sei.cmu.edu/library/considerations-for-evaluating-large-language-models-for-cybersecurity-tasks/.
@techreport{gennari_2024,
author={Gennari, Jeff and Lau, Shing-hon and Perl, Samuel and Parish, Joel and Sastry, Girish},
title={Considerations for Evaluating Large Language Models for Cybersecurity Tasks},
month={Feb},
year={2024},
institution={Software Engineering Institute, Carnegie Mellon University},
url={https://www.sei.cmu.edu/library/considerations-for-evaluating-large-language-models-for-cybersecurity-tasks/},
note={Accessed: 2026-Aug-19}
}
Gennari, Jeff, Shing-hon Lau, Samuel Perl, Joel Parish, and Girish Sastry. "Considerations for Evaluating Large Language Models for Cybersecurity Tasks." Software Engineering Institute, Carnegie Mellon University. Software Engineering Institute, February 20, 2024. https://www.sei.cmu.edu/library/considerations-for-evaluating-large-language-models-for-cybersecurity-tasks/.
J. Gennari, S. Lau, S. Perl, J. Parish, and G. Sastry, "Considerations for Evaluating Large Language Models for Cybersecurity Tasks," Software Engineering Institute, Carnegie Mellon University. Software Engineering Institute, 20-Feb-2024 [Online]. Available: https://www.sei.cmu.edu/library/considerations-for-evaluating-large-language-models-for-cybersecurity-tasks/. [Accessed: 19-Aug-2026].
Gennari, Jeff, Shing-hon Lau, Samuel Perl, Joel Parish, and Girish Sastry. "Considerations for Evaluating Large Language Models for Cybersecurity Tasks." Software Engineering Institute, Carnegie Mellon University, Software Engineering Institute, 20 Feb. 2024. https://www.sei.cmu.edu/library/considerations-for-evaluating-large-language-models-for-cybersecurity-tasks/. Accessed 19 Aug. 2026.
Gennari, Jeff; Lau, Shing-hon; Perl, Samuel; Parish, Joel; & Sastry, Girish. Considerations for Evaluating Large Language Models for Cybersecurity Tasks. Software Engineering Institute. 2024. https://www.sei.cmu.edu/library/considerations-for-evaluating-large-language-models-for-cybersecurity-tasks/