IP Library › Granted Patent US 12,542,795
Granted Patent B2
US 12,542,795 · App. 18/432,940 · Granted Feb 3, 2026

Ai-driven multi-faceted cyber threat classification and categorization

Inventors: Urjitkumar Patel (Scotch Plains, NJ); Chinmay Gondhalekar (Jersey City, NJ); Fang-Chun Yeh (New York, NY); Cristina Polizu (Great Neck, NY)
Assignee: S&P Global Inc.
H04L63/1425G06F18/2415H04L63/1416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,542,795
App. No.
18/432,940
Granted
Feb 3, 2026
Kind
B2
Abstract

Classifying cybersecurity signals from media sources into distinct categories is provided. The method comprises receiving a first data subset comprising data points labeled by subject matter experts according to a predetermined number of specified categories. The data points include information regarding cybersecurity from a set of news articles. The first subset is enriched by applying a random forest algorithm to generate synthetic data points, thereby deriving a second data subset that is augmented from the first subset. The combined first and second data subsets comprise an enhanced training dataset. A BERT model is trained with the enhanced training dataset to classify cybersecurity-related news according to the specified categories. The BERT model utilizes a specialized vector database integrating domain-specific cyber-related terminology and contextual embeddings. The trained BERT model classifies a second set of news articles according to the specified categories. The classification accounts for evolving cybersecurity terminologies and threat landscapes.

Claims (69)

1 . A computer-implemented method for classifying cybersecurity signals from media sources into distinct categories, the method comprising:

receiving a first data subset comprising data points labeled by subject matter experts according to a predetermined number of specified categories, wherein the data points include information regarding cybersecurity from a first set of news articles;

enriching the first data subset by applying a random forest algorithm to generate synthetic data points, thereby deriving a second data subset that is augmented from the first data subset, wherein the combined first and second data subsets comprise an enhanced training dataset;

training a bidirectional encoder representations from transformers (BERT) model with the enhanced training dataset to classify cybersecurity-related news according to the number of specified categories, wherein the BERT model utilizes a specialized vector database integrating domain-specific cyber-related base terminology and contextual embeddings; and

classifying, by the BERT model, a second set of news articles according to the specified categories, wherein the classification accounts for evolving cybersecurity terminologies and threat landscapes.

2 . The method of claim 1 , further comprising:

dynamically updating the vector database by periodically scanning a plurality of media sources to discover new cyber-related terms, wherein the new cyber-related terms satisfy domain-specific acceptance criteria in relation to the base terminology;

adding new cyber-related vectors that represent the new cyber-related terms to the vector database after receiving human confirmation of the new cyber-related terms; and

fetching a third set of news articles according to the new cyber-related terms.

3 . The method of claim 2 , wherein the new cyber-related vectors have a cosine similarity score with an existing cyber term greater than a specified threshold.

4 . The method of claim 2 , wherein the new cyber-related terms neither duplicate nor extend the cyber-related terms already present in the base terminology.

5 . The method of claim 1 , further comprising fine-tuning the BERT model using parameter-efficient fine-tuning.

6 . The method of claim 5 , wherein fine-tuning the BERT model further comprises at least one of:

applying gradient freezing to the BERT model during fine-tuning, wherein the gradient freezing is applied to all layers except layers undergoing fine-tuning;

fine-tuning the last layer of the BERT model in conjunction with a classifier layer; or

fine-tuning the last two layers of the BERT model in conjunction with a classifier layer.

7 . The method of claim 5 , wherein using the parameter-efficient fine-tuning further comprises using low rank adaptation that creates adapter weight matrices which are used with complete model weights to perform classification.

8 . The method of claim 1 , wherein the specified categories of cybersecurity-related news comprise:

recent cyber attack;

cyber-related litigation;

future cyber threats; and

cyber risk prevention.

9 . A system for classifying cybersecurity signals from media sources into distinct categories, the system comprising:

a storage device that stores program instructions;

one or more processors operably connected to the storage device and configured to execute the program instructions to cause the system to:

receive a first data subset comprising data points labeled by subject matter experts according to a predetermined number of specified categories, wherein the data points include information regarding cybersecurity from a first set of news articles;

enrich the first data subset by applying a random forest algorithm to generate synthetic data points, thereby deriving a second data subset that is augmented from the first data subset, wherein the combined first and second data subsets comprise an enhanced training dataset;

train a bidirectional encoder representations from transformers (BERT) model with the enhanced training dataset to classify cybersecurity-related news according to the number of specified categories, wherein the BERT model utilizes a specialized vector database integrating domain-specific cyber-related base terminology and contextual embeddings; and

classify, by the BERT model, a second set of news articles according to the specified categories, wherein the classification accounts for evolving cybersecurity terminologies and threat landscapes.

10 . The system of claim 9 , wherein the processors further execute program instructions for:

dynamically updating the vector database by periodically scanning a plurality of media sources to discover new cyber-related terms, wherein the new cyber-related terms satisfy domain-specific acceptance criteria in relation to the base terminology;

adding new cyber-related vectors that represent new cyber-related terms to the vector database after receiving human confirmation of the new cyber-related terms; and

fetching a third set of news articles according to the new cyber-related terms.

11 . The system of claim 10 , wherein the new cyber-related vectors have a cosine similarity score with an existing cyber term greater than a specified threshold.

12 . The system of claim 10 , wherein the new cyber-related terms neither duplicate nor extend the cyber-related terms already present in the base terminology.

13 . The system of claim 9 , wherein the processors further execute program instructions for fine-tuning the BERT model using parameter-efficient fine-tuning.

14 . The system of claim 13 , wherein fine-tuning the BERT model further comprises at least one of:

applying gradient freezing to the BERT model during fine-tuning, wherein the gradient freezing is applied to all layers except layers undergoing fine-tuning;

fine-tuning the last layer of the BERT model in conjunction with a classifier layer; or

fine-tuning the last two layers of the BERT model in conjunction with a classifier layer.

15 . The system of claim 13 , wherein using the parameter-efficient fine-tuning further comprises using low rank adaptation that creates adapter weight matrices which are used with complete model weights to perform classification.

16 . The system of claim 9 , wherein the specified categories of cybersecurity-related news comprise:

recent cyber attack;

cyber-related litigation;

future cyber threats; and

cyber risk prevention.

17 . A computer program product for classifying cybersecurity signals from media sources into distinct categories, the computer program product comprising:

a computer-readable storage medium having program instructions embodied thereon to perform the steps of:

receiving a first data subset comprising data points labeled by subject matter experts according to a predetermined number of specified categories, wherein the data points include information regarding cybersecurity from a first set of news articles;

enriching the first data subset by applying a random forest algorithm to generate synthetic data points, thereby deriving a second data subset that is augmented from the first data subset, wherein the combined first and second data subsets comprise an enhanced training dataset;

training a bidirectional encoder representations from transformers (BERT) model with the enhanced training dataset to classify cybersecurity-related news according to the number of specified categories, wherein the BERT model utilizes a specialized vector database integrating domain-specific cyber-related base terminology and contextual embeddings; and

classifying, by the BERT model, a second set of news articles according to the specified categories, wherein the classification accounts for evolving cybersecurity terminologies and threat landscapes.

18 . The computer program product of claim 17 , further comprising instructions for:

dynamically updating the vector database by periodically scanning a plurality of media sources to discover new cyber-related terms, wherein the new cyber-related terms satisfy domain-specific acceptance criteria in relation to the base terminology;

adding new cyber-related vectors that represent new cyber-related terms to the vector database after receiving human confirmation of the new cyber-related terms; and

fetching a third set of news articles according to the new cyber-related terms.

19 . The computer program product of claim 18 , wherein the new cyber-related vectors have a cosine similarity score with an existing cyber term greater than a specified threshold.

20 . The computer program product of claim 18 , wherein the new cyber-related terms neither duplicate nor extend the cyber-related terms already present in the base terminology.

21 . The computer program product of claim 17 , further comprising instructions for fine-tuning the BERT model using parameter-efficient fine-tuning.

22 . The computer program product of claim 21 , wherein fine-tuning the BERT model further comprises instructions for at least one of:

applying gradient freezing to the BERT model during fine-tuning, wherein the gradient freezing is applied to all layers except layers undergoing fine-tuning;

fine-tuning the last layer of the BERT model in conjunction with a classifier layer; or

fine-tuning the last two layers of the BERT model in conjunction with a classifier layer.

23 . The computer program product of claim 21 , wherein using the parameter-efficient fine-tuning further comprises instructions for using low rank adaptation that creates adapter weight matrices which are used with complete model weights to perform classification.

24 . The computer program product of claim 17 , wherein the specified categories of cybersecurity-related news comprise:

recent cyber attack;

cyber-related litigation;

future cyber threats; and

cyber risk prevention.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2024
From: PATEL, URJITKUMAR; GONDHALEKAR, CHINMAY; YEH, FANG-CHUN; POLIZU, CRISTINA
To: S&P GLOBAL INC.
Reel/Frame 066399/0705 →
Continuity (1)
Related Publication 20250254187A1 · Aug 7, 2025
References Cited (19)
US 11288364B1 · Savir · 2022 [cited by examiner]
US 12192364B1 · Mullaney · 2025 [cited by examiner]
US 20190363925A1 · Davis · 2019 [cited by examiner]
US 20210073377A1 · Coull · 2021 [cited by examiner]
US 20210200877A1 · Salo · 2021 [cited by examiner]
US 20220292189A1 · Silberman · 2022 [cited by examiner]
US 20220350884A1 · Hencinski · 2022 [cited by examiner]
US 20230195828A1 · Huang · 2023 [cited by examiner]
US 20230385548A1 · Tully · 2023 [cited by examiner]
US 20230412627A1 · Szilágyi · 2023 [cited by examiner]
US 20240275817A1 · Grout · 2024 [cited by examiner]
US 20240354503A1 · Baruch · 2024 [cited by examiner]
US 20240388602A1 · Angiolelli · 2024 [cited by examiner]
US 20240403428A1 · Lal · 2024 [cited by examiner]
US 20250053587A1 · Coulter · 2025 [cited by examiner]
US 20250148472A1 · Ur · 2025 [cited by examiner]
US 20250165616A1 · Cameron · 2025 [cited by examiner]
US 20250209156A1 · Sankaran · 2025 [cited by examiner]
US 20250238510A1 · Sharpe · 2025 [cited by examiner]