IP Library › Granted Patent US 12,205,037
Granted Patent B2
US 12,205,037 · App. 17/081,592 · Granted Jan 21, 2025

Clustering autoencoder

Inventor: Philip A. Sallee (South Riding, VA)
Assignee: Raytheon Company
G06N3/084G06F18/22G06F18/24G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,037
App. No.
17/081,592
Granted
Jan 21, 2025
Kind
B2
Abstract

Discussed herein are devices, systems, and methods for classification using a clustering autoencoder. A method can include receiving, by an encoder of an autoencoder, content, the autoencoder trained using other content and corresponding labels, providing, by the encoder, a latent feature representation of the content to a decoder of the autoencoder, providing, by a clustering layer situated between the encoder and the decoder, a probability that the content belongs to a class of classes represented by respective clusters in a latent feature representation space based on a distance between the feature representation and the cluster, and providing, by the decoder, reconstructed content that is a construction of the content based on the latent feature representation.

Claims (44)

1. A computer-implemented method for content classification using supervised machine learning (ML), the method comprising:

receiving, by an encoder of an autoencoder, content, the autoencoder trained using other content and corresponding labels, the autoencoder including an encoder, decoder, and a clustering layer, the clustering layer situated between the encoder and the decoder;

providing, by the encoder, a latent feature representation in a latent feature space of the content to the decoder and the clustering layer;

determining, by the clustering layer, respective confidences based on respective distances between the latent feature representation and respective representative points in respective clusters, the respective confidences indicating respective probabilities that the content belongs to the respective clusters, each of the respective clusters representing a different respective class of respective classes in the latent feature space;

providing, by the clustering layer, the respective confidences and corresponding classes of the respective classes; and

providing, by the decoder and based on the latent feature representation, reconstructed content that is a construction of the content.

2. The computer-implemented method of claim 1 , wherein the autoencoder is trained to reduce a difference between the content and the reconstructed content and increase a similarity between the labels and predicted class probabilities.

3. The computer-implemented method of claim 2 , wherein the autoencoder is trained based on an L2 norm and a cross-entropy.

4. The computer-implemented method of claim 1 , wherein the probability that the content belongs to the class is determined using a Student's t-Distribution or a mixture of Gaussian distributions.

5. The computer-implemented method of claim 1 , further comprising sampling a latent feature representation from a cluster of the clusters to determine another member of the class.

6. The computer-implemented method of claim 5 , further comprising:

wherein the probability that the content belongs to the class is determined using a Student's t-Distribution;

fitting a mixture of Gaussian distributions to the Student's t-Distribution; and

sampling a latent feature representation from a cluster of the clusters using the mixture of Gaussian distributions to determine another member of the class.

7. The computer-implemented method of claim 1 , further comprising:

determining a reconstruction error based on the reconstructed and the content; and

determining the content is not a member of the class in response to determining the reconstruction error is greater than a threshold.

8. The computer-implemented method of claim 7 , wherein the content is computer-transmissible content and the classes include respective malicious classes.

9. The computer-implemented method of claim 8 , further comprising determining a normalized probability complement plus the reconstruction error to determine the probability that the latent feature representation is a member of the class.

10. The computer-implemented method of claim 1 , further comprising determining the content is an outlier in response to the distance between the latent feature representation of the content being more than a threshold distance away from a central point of each of the clusters.

11. A non-transitory machine-readable medium including instructions that, when executed by a machine, cause the machine to perform operations for content classification using machine learning (ML), the operations comprising:

receiving, by an encoder of an autoencoder, content, the autoencoder trained using other content and corresponding labels, the autoencoder including an encoder, decoder, and a clustering layer, the clustering layer situated between the encoder and the decoder;

providing, by the encoder, a latent feature representation in a latent feature space of the content to the decoder and the clustering layer;

determining, by the clustering layer, respective confidences based on respective distances between the latent feature representation and respective representative points in respective clusters, the respective confidences indicating respective probabilities that the content belongs to the respective clusters, each of the respective clusters representing a different respective class of respective classes in the latent feature space;

providing, by the clustering layer, the respective confidences and corresponding classes of the respective classes; and

providing, by the decoder and based on the latent feature representation, reconstructed content that is a construction of the content.

12. The non-transitory machine-readable medium of claim 11 , wherein the autoencoder is trained to reduce a difference between the content and the reconstructed content and increase a similarity between the labels and predicted class probabilities.

13. The non-transitory machine-readable medium of claim 12 , wherein the autoencoder is trained based on an L2 norm and a cross-entropy.

14. The non-transitory machine-readable medium of claim 11 , wherein the probability that the content belongs to the class is determined using a Student's t-Distribution or a mixture of Gaussian distributions.

15. The non-transitory machine-readable medium of claim 11 , wherein the operations further comprise sampling a latent feature representation from a cluster of the clusters to determine another member of the class.

16. An autoencoder clustering system comprising:

a memory including instructions stored thereon;

processing circuitry configured to execute the instructions, the instruction, when executed by the processing circuitry cause the processing circuitry to implement the clustering autoencoder that:

receives, by an encoder of the autoencoder, content, the autoencoder trained using other content and corresponding labels, the autoencoder including an encoder, decoder, and a clustering layer situated between the encoder and the decoder, the autoencoder including an encoder, decoder, and a clustering layer, the clustering layer situated between the encoder and the decoder;

provides, by the encoder, a latent feature representation in a latent feature space of the content to the decoder;

determines, by the clustering layer, respective confidences based on respective distances between the latent feature representation and respective representative points in respective clusters, the respective confidences indicating respective probabilities that the content belongs to the respective clusters, each of the respective clusters representing a different respective class of respective classes in the latent feature space;

provides, by the clustering layer, the respective confidences and corresponding classes of the respective classes; and

provides, by the decoder and based on the latent feature representation, reconstructed content that is a construction of the content.

17. The system of claim 16 , wherein the instructions, when executed by the processing circuitry, further cause the processing circuitry to:

determine a reconstruction error based on the reconstructed and the content; and

determine the content is not a member of the class in response to determining the reconstruction error is greater than a threshold.

18. The system of claim 17 , wherein the content is computer-transmissible content and the classes include respective malicious classes.

19. The system of claim 18 , wherein the instructions, when executed by the processing circuitry, further cause the processing circuitry to determine a normalized probability complement plus the reconstruction error to determine the probability that the latent feature representation is a member of the class.

20. The of claim 16 , wherein the instructions, when executed by the processing circuitry, further cause the processing circuitry to determine the content is an outlier in response to the distance between the latent feature representation of the content being more than a threshold distance away from a central point of each of the clusters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2020
From: SALLEE, PHILIP A.
To: RAYTHEON COMPANY
Reel/Frame 054183/0927 →
Continuity (1)
Related Publication 20220129758A1 · Apr 28, 2022
References Cited (31)
US 11210537B2 · Koivisto · 2021 [cited by examiner]
US 20180018535A1 · Li et al. · 2018 [cited by applicant]
US 20190188562A1 · Edwards · 2019 [cited by examiner]
US 20190258878A1 · Koivisto · 2019 [cited by examiner]
US 20190347277A1 · Tao et al. · 2019 [cited by applicant]
US 20200104648A1 · Yadav · 2020 [cited by applicant]
US 20210264173A1 · Wang · 2021 [cited by examiner]
US 20210358626A1 · Nicula · 2021 [cited by examiner]
US 20220028180A1 · Wong · 2022 [cited by examiner]
US 20220129706A1 · Vivona · 2022 [cited by examiner]
US 20220129712A1 · Sallee · 2022 [cited by examiner]
US 20220335305A1 · Baker · 2022 [cited by examiner]
US 20230038256A1 · Tal · 2023 [cited by examiner]
US 20230306267A1 · Jacob Banville · 2023 [cited by examiner]
Athalye, Anish, et al., “Synthesizing Robust Adversarial Examples”, International Conference on Machine Learning, (2018), 10 pgs. [cited by applicant]
Carlini, Nicholas, et al., “Defensive Distillation is Not Robust to Adversarial Examples”, arXiv:1607.04311v1, (7/1416), 3 pgs. [cited by applicant]
Evtimov, Ivan, et al., “Robust Physical-World Attacks on Machine Learning Models”, arXiv:1707.08945v2 [cs.CR], (Jul. 30, 2017), 10 pgs. [cited by applicant]
Goodfellow, Ian, et al., “Explaining and Harnessing Adversarial Examples”, Published as a conference paper at ICLR. arXiv:1412.6572v3, (Mar. 20, 2015), 11 pgs. [cited by applicant]
Guo, Chuan, et al., “On Calibration of Modern Neural Networks”, arXiv:1706.04599v2 [cs.LG], (Aug. 3, 2017), 14 pgs. [cited by applicant]
Hein, Matthias, et al., “Why ReLU networks yield high-con?dence predictions far away from the training data and how to mitigate the problem”, IEEE Conference on Computer Vision and Pattern Recognition, (2019), 41-50. [cited by applicant]
Jang, Uyeong, et al., “Objective Metrics and Gradient Descent Algorithms for Adversarial Examples in Machine Learning”, ACSAC, (2017), 262-277. [cited by applicant]
Lee, Kimin, et al., “A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacks”, 32nd Conference on Neural Information Processing Systems, (2018), 11 pgs. [cited by applicant]
Moosavi-Dezfooli, Seyed-Mohsen, et al., “DeepFool: a simple and accurate method to fool deep neural networks”, Computer Vision Foundation, (2015), 2574-2582. [cited by applicant]
Nguyen, Anh, et al., “Deep Neural Networks are Easily Fooled: High Confidence Predictions for Unrecognizable Images”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, (2015), 427-436. [cited by applicant]
Sapre, Suchet, et al., “A Robust Comparison of the KDDCup99 and NSL-KDD IoT Network Intrusion Detection Datasets Through Various Machine Learning Algorithms”, arXiv:1912.13204v1, (Dec. 2019), 8 pgs. [cited by applicant]
Song, Chufeng, et al., “Auto-encoder Based Data Clustering”, CIARP, Part I, LNCS 8258, (2013), 117-124. [cited by applicant]
Van Der Maaten, L., et al., “Visualizing Data Using t-SNE”, J. Mach. Learn. Res. 9, (2008), 2579-2605. [cited by applicant]
Xie, Junyuan, et al., “Unsupervised deep embedding for clustering analysis.”, Proceedings of the 33rd International Conference on Machine Learning, (2016), 10 pgs. [cited by applicant]
“U.S. Appl. No. 17/081,612, Non Final Office Action mailed Dec. 15, 2023”, 36 pgs. [cited by applicant]
“U.S. Appl. No. 17/081,612, Examiner Interview Summary mailed Feb. 23, 2024”, 3 pgs. [cited by applicant]
“U.S. Appl. No. 17/081,612, Response filed Mar. 15, 2024 to Non Final Office Action mailed Dec. 15, 2023”, 8 pgs. [cited by applicant]