IP Library › Granted Patent US 11,921,820
Granted Patent B2
US 11,921,820 · App. 17/018,885 · Granted Mar 5, 2024

Real-time minimal vector labeling scheme for supervised machine learning

Inventor: Sameer T. Khanna (Cupertino, CA)
Assignee: Fortinet, Inc.
G06F18/2155G06F18/10G06F18/2113G06F18/2115G06F18/213G06F18/22G06F18/23G06F18/2321G06F18/24137G06F18/2431G06F18/28G06N3/09G06N5/01G06N20/00G06N3/092
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,921,820
App. No.
17/018,885
Granted
Mar 5, 2024
Kind
B2
Abstract

Systems and methods are described for training a machine learning model using intelligently selected multiclass vectors. According to an embodiment, a set of un-labeled feature vectors are received. The set of feature vectors are grouped into clusters within a vector space having fewer dimensions than the first set of feature vectors by applying a homomorphic dimensionality reduction algorithm to the set of feature vectors and performing centroid-based clustering. An optimal set of clusters among the clusters is identified by performing a convex optimization process on the clusters. Vector labeling is minimized by selecting ground truth representative vectors including a representative vector from each cluster of the optimal set of clusters. A set of labeled feature vectors is created based on labels received from an oracle for each of the representative vectors. A machine-learning model is trained for multiclass classification based on the set of labeled feature vectors.

Claims (99)

1. A method comprising:

receiving, by a processing resource of a computing system, a first set of feature vectors, wherein the first set of feature vectors are un-labeled;

grouping, by the processing resource, the first set of feature vectors into a plurality of clusters within a vector space having fewer dimensions than the first set of feature vectors by applying a homomorphic dimensionality reduction algorithm to the first set of feature vectors and performing centroid-based clustering;

identifying, by the processing resource, an optimal set of clusters among the plurality of clusters by performing a convex optimization process on the plurality of clusters;

minimizing, by the processing resource, vector labeling by selecting a plurality of ground truth representative vectors including a representative vector from each cluster of the optimal set of clusters;

creating, by the processing resource, a set of labeled feature vectors based on labels received from an oracle for each of the plurality of representative vectors;

training, by the processing resource, a machine-learning model for multiclass classification based on the set of labeled feature vectors; and

training the machine-learning model with inductive learning, wherein the inductive learning comprises:

selecting an unlabeled feature vector from the first set of feature vectors;

classifying the un-labeled feature vector using the machine learning model to get a model classified cluster with a confidence score;

determining whether the confidence score is greater than a threshold; and

when said determining is affirmative:

determining a Mahalanobis distance of the un-labeled feature vector with respect to each labeled feature vector of the first set of feature vectors;

determining a statistically matching cluster of labeled feature vectors to which the un-labeled feature vector is closest based on the determined Mahalanobis distance;

determining whether the model classified cluster and the statistically matching cluster are the same; and

when the model classified cluster and the statistically matching cluster are determined to be the same:

labeling the un-labeled feature vector based on the label associated with the model classified cluster; and

model fitting the machine learning model based on the labeling.

2. The method of claim 1 , wherein the homomorphic dimensionality reduction algorithm comprises T-Distributed Stochastic Neighbor Embedding (t-SNE).

3. The method of claim 1 , wherein the centroid-based clustering is based on constructed probability distributions of Cartesian distances between different vectors within the homomorphically translated set.

4. The method of claim 1 , further comprising:

calculating, by the processing resource, a prediction skepticism score for each feature vector of the first set of feature vectors when classified using the machine-learning model based on a skepticism heuristic function; and

selecting, by the processing resource, a boundary condition vector for labeling from the first set of feature vectors, wherein the prediction skepticism score of the boundary condition vector has a highest degree of skepticism.

5. The method of claim 4 , further comprising:

associating, by the processing resource, a label received from the oracle with the boundary condition vector; and

retraining, by the processing resource, the machine-learning model.

6. The method of claim 1 , wherein the plurality representative vectors are selected based on a distance from the center of their respective clusters of the optimal set of clusters.

7. The method of claim 1 , further comprising inductively forgetting a feature vector; wherein said inductively forgetting comprises:

selecting a labeled feature vector from a set of feature vectors that have been labeled through inductive learning;

classifying the labeled feature vector using the machine learning model to get a model classified cluster with a confidence score;

determining whether the confidence score is lower than a base threshold; and

when said determining is affirmative:

determining a Mahalanobis distance of the labeled feature vector with respect to other labeled feature vector of the first set of feature vectors;

determining a statistically matching cluster of labeled feature vectors to which the labeled feature vector is closest based on the determined Mahalanobis distance;

determining whether the model classified cluster and the statistically matching cluster are the same; and

when the model classified cluster and the statistically matching cluster are not the same:

un-labelling the labeled feature vector; and

model fitting the machine learning model based on the un-labeling.

8. A system comprising:

a processing resource; and

a non-transitory computer-readable medium, coupled to the processing resource, having stored therein instructions that when executed by the processing resource cause the processing resource to:

receive a first set of feature vectors, wherein the first set of feature vectors are un-labeled;

group the first set of feature vectors into a plurality of clusters within a vector space having fewer dimensions than the first set of feature vectors by applying a homomorphic dimensionality reduction algorithm to the first set of feature vectors and performing centroid-based clustering;

identify an optimal set of clusters among the plurality of clusters by performing a convex optimization process on the plurality of clusters;

minimize vector labeling by selecting a plurality of ground truth representative vectors including a representative vector from each cluster of the optimal set of clusters;

create a set of labeled feature vectors based on labels received from an oracle for each of the plurality of representative vectors;

train a machine-learning model for multiclass classification based on the set of labeled feature vectors; and

train the machine-learning model with inductive learning, wherein the inductive learning comprises:

selecting an unlabeled feature vector from the first set of feature vectors;

classifying the un-labeled feature vector using the machine learning model to get a model classified cluster with a confidence score;

determining whether the confidence score is greater than a threshold; and

when said determining is affirmative:

determining a Mahalanobis distance of the un-labeled feature vector with respect to each labeled feature vector of the first set of feature vectors;

determining a statistically matching cluster of labeled feature vectors to which the un-labeled feature vector is closest based on the determined Mahalanobis distance;

determining whether the model classified cluster and the statistically matching cluster are the same; and

when the model classified cluster and the statistically matching cluster are determined to be the same:

labeling the un-labeled feature vector based on the label associated with the model classified cluster; and

model fitting the machine learning model based on the labeling.

9. The system of claim 8 , wherein the homomorphic dimensionality reduction algorithm comprises T-Distributed Stochastic Neighbor Embedding (t-SNE).

10. The system of claim 8 , wherein the centroid-based clustering is based on constructed probability distributions of Cartesian distances between different vectors within the homomorphically translated set.

11. The system of claim 8 , wherein the instructions further cause the processing resource to:

calculate a prediction skepticism score for each feature vector of the first set of feature vectors when classified using the machine-learning model based on a skepticism heuristic function; and

select a boundary condition vector for labeling from the first set of feature vectors, wherein the prediction skepticism score of the boundary condition vector has a highest degree of skepticism.

12. The system of claim 11 , wherein the instructions further cause the processing resource to:

associate a label received from the oracle with the boundary condition vector; and

retrain the machine-learning model.

13. The system of claim 8 , wherein the plurality representative vectors are selected based on a distance from the center of their respective clusters of the optimal set of clusters.

14. The system of claim 8 , wherein the instructions further cause the processing resource to inductively forget a feature vector, wherein inductively forgetting comprises:

selecting a labeled feature vector from a set of feature vectors that have been labeled through inductive learning;

classifying the labeled feature vector using the machine learning model to get a model classified cluster with a confidence score;

determining whether the confidence score is lower than a base threshold; and

when said determining is affirmative:

determining a Mahalanobis distance of the labeled feature vector with respect to other labeled feature vector of the first set of feature vectors;

determining a statistically matching cluster of labeled feature vectors to which the labeled feature vector is closest based on the determined Mahalanobis distance;

determining whether the model classified cluster and the statistically matching cluster are the same; and

when the model classified cluster and the statistically matching cluster are not the same:

un-labelling the labeled feature vector; and

model fitting the machine learning model based on the un-labeling.

15. The system of claim 8 , wherein the first set of feature vectors are representative of a plurality of types of Internet of Things (IoT) devices and the machine learning model is trained for classifying IoT devices, and wherein the machine learning model is deployed on a network security device.

16. A non-transitory machine readable medium storing instructions that when executed by a processing resource of a computer system cause the processing resource to:

receive a first set of feature vectors, wherein the first set of feature vectors are un-labeled;

group the first set of feature vectors into a plurality of clusters within a vector space having fewer dimensions than the first set of feature vectors by applying a homomorphic dimensionality reduction algorithm to the first set of feature vectors and performing centroid-based clustering;

identify an optimal set of clusters among the plurality of clusters by performing a convex optimization process on the plurality of clusters;

minimize vector labeling by selecting a plurality of ground truth representative vectors including a representative vector from each cluster of the optimal set of clusters;

create a set of labeled feature vectors based on labels received from an oracle for each of the plurality of representative vectors;

train a machine-learning model for multiclass classification based on the set of labeled feature vectors; and

train the machine-learning model with inductive learning, wherein the inductive learning comprises:

selecting an unlabeled feature vector from the first set of feature vectors;

classifying the un-labeled feature vector using the machine learning model to get a model classified cluster with a confidence score;

determining whether the confidence score is greater than a threshold; and

when said determining is affirmative:

determining a Mahalanobis distance of the un-labeled feature vector with respect to each labeled feature vector of the first set of feature vectors;

determining a statistically matching cluster of labeled feature vectors to which the un-labeled feature vector is closest based on the determined Mahalanobis distance;

determining whether the model classified cluster and the statistically matching cluster are the same; and

when the model classified cluster and the statistically matching cluster are determined to be the same:

labeling the un-labeled feature vector based on the label associated with the model classified cluster; and

model fitting the machine learning model based on the labeling.

17. The non-transitory machine readable medium of claim 16 , wherein the homomorphic dimensionality reduction algorithm comprises T-Distributed Stochastic Neighbor Embedding (t-SNE).

18. The non-transitory machine readable medium of claim 16 , wherein the centroid-based clustering is based on constructed probability distributions of Cartesian distances between different vectors within the homomorphically translated set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2020
From: KHANNA, SAMEER T.
To: FORTINET, INC.
Reel/Frame 053751/0083 →
Continuity (1)
Related Publication 20220083815A1 · Mar 17, 2022
Cited By (2)
US 12,335,116 US 12,482,242