IP Library › Granted Patent US 12,572,441
Granted Patent B2
US 12,572,441 · App. 17/812,453 · Granted Mar 10, 2026

Fully unsupervised pipeline for clustering anomalies detected in computerized systems

Inventors: Ioana Giurgiu (Zürich, CH); Artur Dox (Hofheim, DE); Mircea R. Gusat (Langnau am Albis, CH)
Assignee: International Business Machines Corporation
G06F11/3419G06F11/076G06F18/2321G06N3/0455G06N3/0464
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,441
App. No.
17/812,453
Filed
Jul 14, 2022
Granted
Mar 10, 2026
Kind
B2
Art Unit
2128
USPC
706/15
Abstract

The invention is notably directed to a computer-implemented method of clustering anomalies detected in a computerized system. The proposed method makes use of an unsupervised cognitive model, executed based on input datasets to obtain clusters of anomalies. The method accesses input datasets, which correspond to detected anomalies of the computerized system. These anomalies span respective time windows. Each input dataset comprises a set of timeseries of key performance indicators. The key performance indicators of each input dataset extend over a respective time window. That is, each anomaly corresponds to a respective time window. This model includes a first stage, which includes an encoder designed to learn fixed-size representations of input datasets, and a second stage, which is a clustering stage. The model is executed based on the input datasets accessed, the first stage learning fixed-size representations of the input datasets and the second stage clustering the learned representations.

Claims (57)

1 . A computer-implemented method of clustering anomalies detected in a computerized system, wherein the method comprises:

accessing input datasets corresponding to detected anomalies of the computerized system, wherein the anomalies span respective time windows and each of the corresponding input datasets comprises a set of timeseries of key performance indicators extending over a respective one of the time windows;

loading an unsupervised cognitive model, wherein the model includes a first stage and a second stage, wherein the first stage includes an encoder designed to learn fixed-size representations of given datasets and the second stage is a clustering stage, wherein the unsupervised cognitive model is trained in an iterative and alternate manner, based on a combined loss functions, using an early stopping strategy, wherein the model is established by combining a stacked dilated causal convolutional neural network (CNN) encoder-only architecture with a composed loss function with cosine similarity-based negative sampling and the iterative training; and

executing the unsupervised cognitive model based on the input datasets accessed for the first stage to learn fixed-size representations of the input datasets and the second stage to cluster the learned representations, to obtain clusters of anomalies.

2 . The method according to claim 1 , wherein:

the unsupervised cognitive model is executed using a composed loss function combining a first loss function and a second loss function; and

the first loss function and the second loss function are respectively designed for optimizing the representations and the clusters.

3 . The method according to claim 2 , wherein:

the first loss function is designed as a triplet loss function ensuring that the representation learned for a reference portion of each dataset of the input datasets is, on average, closer to the representations learned for distinct portions of said each dataset than to the representations learned for other portions of other ones of the input datasets, wherein each of the reference portion, the distinct portions, and the other portions, corresponds to a respective time segment.

4 . The method according to claim 2 , wherein:

the unsupervised cognitive model is executed by alternately executing the first stage and the second stage, iteratively, with an objective to decrease the composed loss function, such that the representations learned, and the clusters obtained are alternately optimized through several iterations.

5 . The method according to claim 4 , wherein:

the unsupervised cognitive model is executed so as to achieve predefined structural properties of the clusters.

6 . The method according to claim 5 , wherein:

the predefined structural properties are achieved based on silhouette scores of the clusters.

7 . The method according to claim 6 , wherein each of said iterations comprises:

at the first stage:

obtaining representations of the input datasets; and

computing a first loss with the first loss function, based on the representations obtained;

at the second stage:

running a k-means algorithm based on the previously obtained representations;

optimizing a number of clusters based on the silhouette scores obtained for the clusters;

computing a second loss with the second loss function; and

computing a current composed loss based on the first loss and the second loss; and

deciding whether to stop the training by comparing the current composed loss with a previous composed loss, as obtained during a previous one of the iterations.

8 . The method according to claim 1 , wherein:

the encoder is configured as an exponentially dilated, causal convolutional neural network.

9 . The method according to claim 8 , wherein:

the unsupervised cognitive model is designed as a single network, and

the second stage is implemented by outer neural layers of the cognitive model, the outer neural layers connected in output of the first stage.

10 . The method according to claim 9 , wherein:

the first stage includes at least two convolution blocks; and

each of the at least two convolution blocks comprises one or more dilated causal convolutional layers.

11 . The method according to claim 10 , wherein:

the first stage further includes a hierarchy of neural layers arranged in output of each of the dilated causal convolutional layers.

12 . The method according to claim 11 , wherein:

the hierarchy of neural layers includes a weight normalization layer and an activation layer.

13 . The method according to claim 12 , wherein:

the activation layer is a leaky rectified linear unit.

14 . The method according to claim 12 , wherein:

the first stage further comprises a global max pooling layer arranged in an output of the at least two convolution blocks.

15 . The method according to claim 14 , wherein:

the first stage further comprises a linear transformation layer in an output of the global max pooling layer.

16 . The method according to claim 1 , wherein the method further comprises:

identifying types of anomalies corresponding to each of the clusters obtained.

17 . The method according to claim 1 , wherein:

the method further comprises, prior to accessing the input datasets and executing the unsupervised cognitive model, monitoring the computerized system to detect said anomalies.

18 . A computer program for clustering anomalies detected in a computerized system, the computer program product comprising:

one or more computer-readable tangible storage media and program instructions stored on at least one of the one or more tangible storage media, the program instructions executable by a processor capable of performing a method, the method comprising:

accessing input datasets corresponding to detected anomalies of the computerized system, wherein the anomalies span respective time windows and each of the corresponding input datasets comprises a set of timeseries of key performance indicators extending over a respective one of the time windows;

loading an unsupervised cognitive model, wherein the model includes a first stage and a second stage, wherein the first stage includes an encoder designed to learn fixed-size representations of given datasets and the second stage is a clustering stage, wherein the unsupervised cognitive model is trained in an iterative and alternate manner, based on a combined loss functions, using an early stopping strategy, wherein the model is established by combining a stacked dilated causal convolutional neural network (CNN) encoder-only architecture with a composed loss function with cosine similarity-based negative sampling and the iterative training; and

executing the unsupervised cognitive model based on the input datasets accessed for the first stage to learn fixed-size representations of the input datasets and the second stage to cluster the learned representations, to obtain clusters of anomalies.

19 . The computer program according to claim 18 , wherein:

the unsupervised cognitive model involves a composed loss function combining a first loss function and a second loss function, wherein the first loss function and the second loss function are designed for optimizing the representations and the clusters, respectively, and,

in operation, executing the first stage and the second stage, iteratively, to decrease the composed loss function.

20 . The computer program according to claim 18 , wherein:

the first loss function is designed as a triplet loss function ensuring that the representation learned for a reference portion of each dataset of the input datasets is, on average, closer to the representations learned for distinct portions of said each dataset than to the representations learned for other portions of other ones of the input datasets, wherein each of the reference portion, the distinct portions, and the other portions, corresponds to a respective time segment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2022
From: GIURGIU, IOANA; DOX, ARTUR; GUSAT, MIRCEA R.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 060501/0697 →
Continuity (1)
Related Publication 20240020613A1 · Jan 18, 2024
References Cited (18)
US 10643138B2 · Bellala · 2020 [cited by applicant]
US 10904114B2 · Thampy · 2021 [cited by examiner]
US 11636125B1 · Carmona Perez · 2023 [cited by examiner]
US 20160342903A1 · Shumpert · 2016 [cited by applicant]
US 20200097810A1 · Hetherington · 2020 [cited by applicant]
US 20200250812A1 · Ceccaldi · 2020 [cited by examiner]
US 20220019888A1 · Aggarwal · 2022 [cited by examiner]
CN 111914873A · 2020 [cited by applicant]
WO 2021105799A1 · 2021 [cited by applicant]
S. He, B. Yang and Q. Qiao, “Overview of Key Performance Indicator Anomaly Detection,” 2021 IEEE Region 10 Symposium (TENSYMP), Jeju, Korea, Republic of, 2021, pp. 1-6, doi: 10.1 109/TENSYMP 52854.2021.9550989 (Year: 20… [cited by examiner]
“Patent Cooperation Treaty PCT International Search Report”, Applicant's File Reference: P202104754, International Application No. PCT /1B2023/056601, International Filing Date: Jun. 27, 2023, Date of Mailing: Oct. 24, … [cited by applicant]
Chandrakala et al., “A Density based Method for Multivariate Time Series Clustering in Kernel Feature Space,” IEEE, 2008, https://ieeexplore.ieee.org/document/4634055, pp. 1885-1890. [cited by applicant]
Franceschi et al., “Unsupervised Scalable Representation Learning for Multivariate Time Series,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), arXiv:1901.10738v4 [cs.LG] Jan. 3, 2020, https://… [cited by applicant]
Giurgiu, “Efficient Clustering of Multivariate Variable Length Time Series with Representation Learning,” ACM, Conference'17, Jul. 2017, https://dl.acm.org/doi/10.1145/1122445.1122456, 10 pages. [cited by applicant]
Kremer et al., “Mining of Temporal Coherent Subspace Clusters in Multivariate Time Series Databases,” ResearchGate, Conference: Proceedings of the 16th Pacific-Asia conference on Advances in Knowledge Discovery and Data… [cited by applicant]
Li et al., “Clustering-based Anomaly Detection in Multivariate Time Series Data,” Applied Soft Computing Journal, 2020, https://doi.org/10.1016/j.asoc.2020.106919, 37 pages. [cited by applicant]
Ma et al., “Learning Representations for Time Series Clustering,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), https://papers.nips.cc/paper/2019/file/1359aa933b48b754a2f54adb688bfa77-Paper.pd… [cited by applicant]
Zhou et al., “A Model-Based Multivariate Time Series Clustering Algorithm,” PAKDD 2014 Workshops, LNAI 8643, pp. 805-817, 2014, DOI: 10.1007/978-3-319-13186-3_72. [cited by applicant]