IP Library › Granted Patent US 11,743,133
Granted Patent B2
US 11,743,133 · App. 17/474,820 · Granted Aug 29, 2023

Automatic anomaly detection

Inventors: Ke Wei Wei (Beijing, CN); Wei Liu (Beijing, CN); Guo Ran Sun (Beijing, CN); Shuang YS Yu (Beijing, CN); Meichi Maggie Lin (San Jose, CA); Yi Dai (Beijing, CN)
Assignee: International Business Machines Corporation
H04L41/16G06F16/3347G06F16/35G06F18/22G06N20/00H04L41/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,743,133
App. No.
17/474,820
Filed
Sep 14, 2021
Granted
Aug 29, 2023
Kind
B2
Art Unit
2455
USPC
709/223
Abstract

A method includes generating a plurality of vectors representing words in a plurality of documents about an information technology (IT) system and clustering the plurality of vectors to produce a plurality of clusters. The method also includes identifying a cluster of the plurality of clusters that contains a plurality of clustered vectors, generating a feature based on a plurality of words represented by the plurality of clustered vectors, and training a machine learning model to identify an anomaly in the IT system based on the feature.

Claims (40)

1. A method comprising:

generating a plurality of vectors representing words in a plurality of documents about an information technology (IT) system;

clustering the plurality of vectors to produce a plurality of clusters;

generating a feature based on a plurality of words represented by a plurality of clustered vectors in a cluster of the plurality of clusters, wherein the feature indicates a ratio of (i) a number of appearances of a first word of the plurality of words in the plurality of documents and (ii) a number of appearances of a second word of the plurality of words in the plurality of documents;

determining a first metric identified by the first word;

determining a second metric identified by the second word; and

training a machine learning model to identify an anomaly in the IT system while monitoring the first metric and the second metric.

2. The method of claim 1 , further comprising applying the machine learning model to the system to detect the anomaly.

3. The method of claim 1 , wherein the plurality of clustered vectors indicates a relationship between or amongst the plurality of words.

4. The method of claim 1 , wherein a vector of the plurality of vectors indicates at least one of a number of occurrences of a word in the plurality of documents or a proximity of the word to another word in the plurality of documents.

5. The method of claim 1 , wherein the plurality of documents comprises at least one of an instruction manual, a blog, or an issue ticket.

6. The method of claim 1 , further comprising generating, based on a second plurality of words represented by a second plurality of clustered vectors, a second feature, wherein the machine learning model is trained to identify the anomaly further based on the second feature.

7. The method of claim 1 , further comprising excluding a stop word in the plurality of documents from being considered when generating the plurality of vectors, wherein the stop word is identified based on a dictionary that comprises the stop word.

8. The method of claim 1 , wherein generating the feature is further based on a vector in a second cluster, wherein the second cluster has a proximity to the cluster that meets a threshold proximity.

9. The method of claim 1 , further comprising replacing a word in the plurality of documents with a synonym prior to generating the plurality of vectors.

10. An apparatus comprising:

a memory; and

a hardware processor communicatively coupled to the memory, the hardware processor configured to:

generate a plurality of vectors representing words in a plurality of documents about an information technology (IT) system;

cluster the plurality of vectors to produce a plurality of clusters;

generate a feature based on a plurality of words represented by the plurality of clustered vectors, wherein the feature indicates a ratio of (i) a number of appearances of a first word of the plurality of words in the plurality of documents and (ii) a number of appearances of a second word of the plurality of words in the plurality of documents;

determine a first metric identified by the first word;

determine a second metric identified by the second word; and

train a machine learning model to identify an anomaly in the IT system while monitoring the first metric and the second metric.

11. The apparatus of claim 10 , the hardware processor further configured to apply the machine learning model to the system to detect the anomaly.

12. The apparatus of claim 10 , wherein the plurality of clustered vectors indicates a relationship between or amongst the plurality of words.

13. The apparatus of claim 10 , wherein a vector of the plurality of vectors indicates at least one of a number of occurrences of a word in the plurality of documents or a proximity of the word to another word in the plurality of documents.

14. The apparatus of claim 10 , wherein the plurality of documents comprises at least one of an instruction manual, a blog, or an issue ticket.

15. The apparatus of claim 10 , the hardware processor further configured to generate, based on a second plurality of words represented by a second plurality of clustered vectors, a second feature, wherein the machine learning model is trained to identify the anomaly further based on the second feature.

16. The apparatus of claim 10 , the hardware processor further configured to exclude a stop word in the plurality of documents from being considered when generating the plurality of vectors, wherein the stop word is identified based on a dictionary that comprises the stop word.

17. The apparatus of claim 10 , the hardware processor further configured to replace a word in the plurality of documents with a synonym prior to generating the plurality of vectors.

18. A method comprising:

generating a plurality of vectors representing words in a plurality of documents about an information technology (IT) system, wherein the plurality of vectors indicates at least one of a number of occurrences of a word in the plurality of documents or a proximity of the word to another word in the plurality of documents;

clustering the plurality of vectors to produce a plurality of clusters;

generating features based on the plurality of clusters, wherein the feature indicates a ratio of (i) a number of appearances of a first word in the plurality of documents and (ii) a number of appearances of a second word in the plurality of documents;

determining a first metric identified by the first word;

determining a second metric identified by the second word; and

training a machine learning model to identify an anomaly in the IT system while monitoring the first metric and the second metric.

19. The method of claim 18 , further comprising applying the machine learning model to the system to detect the anomaly.

20. The method of claim 18 , wherein a vector of the plurality of vectors indicates at least one of a number of occurrences of a word in the plurality of documents or a proximity of the word to another word in the plurality of documents.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2021
From: WEI, KE WEI; LIU, WEI; SUN, GUO RAN; YU, SHUANG YS; LIN, MEICHI MAGGIE; DAI, YI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 057479/0328 →
Continuity (1)
Related Publication 20230078661A1 · Mar 16, 2023