Threat modeling refers to the proactive and systematic evaluation of threats and risks at the level of software architecture and design and is considered an essential part of secure software engineering. Recent automation efforts, particularly those involving Large-Language Models (llms), show significant promise for improving the overall cost-effectiveness, efficiency, and thus practical feasibility of these activities. Partly because of these developments, and partly due to a growing diversity in threat modeling methods and tools, there is a current lack of clarity on what the targeted qualities of threat modeling efforts should be. In this SoK paper, we address this gap by outlining the different evaluation approaches, criteria, and metrics that have been applied in academic literature to evaluate threat modeling methods, frameworks, and tools. We identify and survey 29 empirical studies with particular attention to their empirical approach. From this synthesis, we construct a quality model for threat modeling, a taxonomy that integrates the diverse quality factors and dimensions into a standalone and hierarchical knowledge structure. The quality model is validated through a comprehensive gap analysis applied to (i) the threat modeling manifesto, and (ii) a validation set of 12 scientific studies published more recently. Based on our overall findings and results, we finally reflect on the future role of llm automation in threat modeling and emphasize the urgent need for published and verified application cases, common empirical benchmarks, and normative threat modeling datasets.

SoK: A quality model for threat modeling methods and tools

Raciti, Mario;Mollaeefar, Majid;Ranise, Silvio;
2026-01-01

Abstract

Threat modeling refers to the proactive and systematic evaluation of threats and risks at the level of software architecture and design and is considered an essential part of secure software engineering. Recent automation efforts, particularly those involving Large-Language Models (llms), show significant promise for improving the overall cost-effectiveness, efficiency, and thus practical feasibility of these activities. Partly because of these developments, and partly due to a growing diversity in threat modeling methods and tools, there is a current lack of clarity on what the targeted qualities of threat modeling efforts should be. In this SoK paper, we address this gap by outlining the different evaluation approaches, criteria, and metrics that have been applied in academic literature to evaluate threat modeling methods, frameworks, and tools. We identify and survey 29 empirical studies with particular attention to their empirical approach. From this synthesis, we construct a quality model for threat modeling, a taxonomy that integrates the diverse quality factors and dimensions into a standalone and hierarchical knowledge structure. The quality model is validated through a comprehensive gap analysis applied to (i) the threat modeling manifesto, and (ii) a validation set of 12 scientific studies published more recently. Based on our overall findings and results, we finally reflect on the future role of llm automation in threat modeling and emphasize the urgent need for published and verified application cases, common empirical benchmarks, and normative threat modeling datasets.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11582/374307
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
social impact