IP Library › Granted Patent US 12,524,539
Granted Patent B2
US 12,524,539 · App. 18/384,954 · Granted Jan 13, 2026

LLM-powered threat modeling

Inventors: Tiferet Ahavah Gazit (Albany, CA); Aditya Sharad (Mountain View, CA)
Assignee: Microsoft Technology Licensing, LLC
G06F21/563G06N3/09G06F2221/033
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,539
App. No.
18/384,954
Granted
Jan 13, 2026
Kind
B2
Abstract

Techniques for implementing an AI threat modeling tool are disclosed. A static analysis tool is used to extract a candidate code snippet from a code repository. The candidate code snippet is identified as potentially being a security relevant code element. The static analysis tool generates additional context associated with the candidate code snippet. An LLM prompt is generated. This prompt is structured to include the candidate code snippet, the context, and a directive to assign a classification to the candidate code snippet. The classification includes a source classification, a sink classification, a sanitizer classification, or a flow step classification. The LLM operates on the prompt to generate output comprising a specific classification for the candidate code snippet. The output is formatted into a data extension file that is consumable by the static analysis tool.

Claims (39)

1 . A method for implementing an artificial intelligence (Al) threat modeling tool, said method being performed by a computing service, wherein the computing service is one of a cloud service operating in a cloud environment, a local service operating on a local device, or a hybrid service that includes a cloud component operating in the cloud environment and a local component operating on the local device, said method comprising:

using a static analysis tool to extract a candidate code snippet from a code repository, wherein the candidate code snippet is identified by the static analysis tool as potentially being a security relevant code element;

using the static analysis tool to generate additional context associated with the candidate code snippet;

generating a large language model (LLM) prompt, which is structured to include the candidate code snippet, the context, and a directive to assign a classification to the candidate code snippet, said classification including a source classification, a sink classification, a sanitizer classification, or a flow step classification, and wherein the directive further includes a directive to determine a type for the classification;

triggering the LLM to operate on the LLM prompt, wherein, as a result of said operating, the LLM generates output comprising a specific classification for the candidate code snippet and a specific type for the specific classification;

formatting the output of the LLM into one or more data extension files that are consumable by the static analysis tool; and

including the one or more data extension files in a corpus of data extension files that are consumable by the static analysis tool.

2 . The method of claim 1 , wherein the context includes one or more few-shot examples.

3 . The method of claim 2 , wherein the one or more few-shot examples include at least one of: a positive example of a sink, a positive example of a source, a positive example of a sanitizer, or a positive example of a flow step.

4 . The method of claim 2 , wherein the one or more few-shot examples include at least one of: a negative example of a sink, a negative example of a source, a negative example of a sanitizer, or a negative example of a flow step.

5 . The method of claim 1 , wherein the context further includes documentation associated with the candidate code snippet.

6 . The method of claim 1 , wherein the context further includes content extracted from a selected number of code lines preceding a code line comprising the candidate code snippet.

7 . The method of claim 1 , wherein the context further includes content extracted from a selected number of code lines succeeding a code line comprising the candidate code snippet.

8 . The method of claim 1 , wherein formatting the output of the LLM into the data extension file includes parsing the output and organizing the parsed output into a format that is consumable by the static analysis tool.

9 . The method of claim 1 , wherein the candidate code snippet is one of a plurality of candidate code snippets that are identified and extracted by the static analysis tool, and wherein a pre-filtering operation is performed on the plurality of candidate code snippets using a defined set of parameters to reduce a number of candidate code snippets included in the plurality of candidate code snippets.

10 . The method of claim 1 , wherein the prompt includes details regarding a particular classification provided to a few-shot example that is also included in the prompt.

11 . A computer system comprising:

a processor system; and

a storage system that includes instructions that are executable by the processor system to cause the computer system to:

use a static analysis tool to extract a candidate code snippet from a code repository, wherein the candidate code snippet is identified by the static analysis tool as potentially being a security relevant code element;

use the static analysis tool to generate additional context associated with the candidate code snippet;

generate a large language model (LLM) prompt, which is structured to include the candidate code snippet, the context, and a directive to assign a classification to the candidate code snippet, said classification including a source classification, a sink classification, a sanitizer classification, or a flow step classification, and wherein the directive further includes a directive to determine a type for the classification;

trigger the LLM to operate on the LLM prompt, wherein, as a result of said operating, the LLM generates output comprising a specific classification for the candidate code snippet and a specific type for the specific classification;

format the output of the LLM into a data extension file that is consumable by the static analysis tool; and

include the data extension file in a corpus of data extension files that are consumable by the static analysis tool.

12 . The computer system of claim 11 , wherein the prompt is generated in real-time.

13 . The computer system of claim 11 , wherein the context includes information obtained from a source that is external to the code repository.

14 . The computer system of claim 11 , wherein the prompt is a fillable, pre-generated template.

15 . The computer system of claim 11 , wherein a size of the prompt is limited by a threshold size.

16 . The computer system of claim 11 , wherein the context includes a plurality of few-shot examples, and wherein a set of the few-shot examples are designed to encourage the LLM to assign negative classification to certain types of candidate code snippets.

17 . The computer system of claim 16 , wherein one type of candidate code snippet is structured to encourage the LLM to classify a particular candidate code snippet as not being a source or a sink.

18 . A method for implementing an artificial intelligence (Al) threat modeling tool, said method being performed by a computing service, wherein the computing service is one of a cloud service operating in a cloud environment, a local service operating on a local device, or a hybrid service that includes a cloud component operating in the cloud environment and a local component operating on the local device, said method comprising:

using a static analysis tool to extract a candidate code snippet from a code repository, wherein the candidate code snippet is identified by the static analysis tool as potentially being a security relevant code element;

using the static analysis tool to generate additional context associated with the candidate code snippet;

generating a large language model (LLM) prompt, which is structured to include the candidate code snippet, the context, and a directive to assign a classification to the candidate code snippet, said classification including a source classification, a sink classification, a sanitizer classification, or a flow step classification;

triggering the LLM to operate on the LLM prompt, wherein, as a result of said operating, the LLM generates output comprising a specific classification for the candidate code snippet; and

formatting the output of the LLM into a data extension file that is consumable by the static analysis tool.

19 . The method of claim 18 , wherein the prompt is a batch prompt that includes multiple candidate code snippets.

20 . The method of claim 18 , wherein the prompt further includes natural language instructions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: GAZIT, TIFERET AHAVAH; SHARAD, ADITYA
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065409/0151 →
Continuity (1)
Related Publication 20250139243A1 · May 1, 2025
References Cited (32)
US 9754112B1 · Moritz · 2017 [cited by examiner]
US 12204644B1 · Cocea · 2025 [cited by examiner]
US 12314690B2 · Ye · 2025 [cited by examiner]
US 20230418815A1 · Zorn · 2023 [cited by examiner]
US 20240028312A1 · Gillman · 2024 [cited by examiner]
US 20240202405A1 · Lang · 2024 [cited by examiner]
US 20240330480A1 · Roytman · 2024 [cited by examiner]
US 20240403438A1 · Chan · 2024 [cited by examiner]
US 20240403545A1 · Nahum · 2024 [cited by examiner]
US 20240411666A1 · Chan · 2024 [cited by examiner]
US 20250028818A1 · Kim · 2025 [cited by examiner]
US 20250028823A1 · Kim · 2025 [cited by examiner]
US 20250028825A1 · Kim · 2025 [cited by examiner]
US 20250028826A1 · Kim · 2025 [cited by examiner]
US 20250028827A1 · Kim · 2025 [cited by examiner]
US 20250030704A1 · Kim · 2025 [cited by examiner]
US 20250036778A1 · Chan · 2025 [cited by examiner]
US 20250097237A1 · Parla · 2025 [cited by examiner]
US 20250117480A1 · Wagh · 2025 [cited by examiner]
US 20250123814A1 · Mcmorran · 2025 [cited by examiner]
US 20250139243A1 · Gazit · 2025 [cited by examiner]
US 20250240313A1 · Zhang · 2025 [cited by examiner]
CN 113779590A · 2021 [cited by applicant]
CN 112835620B · 2022 [cited by applicant]
Ahmed, et al., “Improving Few-shot Prompts with Relevant Static Analysis Products”, Aug. 29, 2023, pp. 1-12. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/049684, Dec. 6, 2024, 15 pages. [cited by applicant]
Liu, et al., “Harnessing the Power of LLM to Support Binary Taint Analysis”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 12, 2023, 12 pages. [cited by applicant]
Pearce, et al., “Can OpenAI Codex and Other Large Language Models Help Us Fix Security Bugs?”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Dec. 3, 2021, 16 pages. [cited by applicant]
Dv, Raja Rao, “Using AI to write secure code with Semgrep”, Retrieved from: https://semgrep.dev/blog/2023/using-ai-to-write-secure-code-with-semgrep/, Apr. 4, 2023, 9 Pages. [cited by applicant]
Lysenko, Mikola, “Introducing Socket AI—ChatGPT-Powered Threat Analysis”, Retrieved from: https://socket.dev/blog/introducing-socket-ai-chatgpt-powered-threat-analysis, Mar. 31, 2023, 7 Pages. [cited by applicant]
Pistoia, et al., “Combining Static Code Analysis and Machine Learning for Automatic Detection of Security Vulnerabilities in Mobile Apps”, In Book Mobile Application Development, Usability, and Security, Jan. 1, 2017. [cited by applicant]
Tiferet Gazit, “github / codeql”, Retrieved from: https://github.com/github/codeql/blob/tiferet/codex/java/ql/experimental/adaptivethreatmodeling/src/ExtractSinkCandidates.ql, Mar. 16, 2023, 2 Pages. [cited by applicant]
Cited By (1)
US 12,591,417