IP Library Patent Application 17115432
Patent Application
App. No. 17/115,432

Inter-Feature Influence in Unlabeled Datasets

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/115,432
Abstract

In one set of embodiments, a computer system can receive an unlabeled dataset comprising a plurality of unlabeled data instances, each unlabeled data instance including values for a plurality of features. The computer system can train, for each feature, a supervised machine learning (ML) model on a labeled dataset derived from the unlabeled dataset, where the labeled dataset comprises a plurality of labeled data instances, and wherein each labeled data instance includes (1) a label corresponding to a value for the feature in an unlabeled data instance of the unlabeled dataset, and (2) values for other features in the unlabeled data instance. The computer system can then compute, for each pair of first and second features in the plurality of features, an inter-feature influence score using the trained supervised ML model for the second feature, the inter-feature influence score indicating how useful the first feature is in predicting the second feature.

Claims (68)

1 . A method comprising:

receiving, by a computer system, an unlabeled dataset comprising a plurality of unlabeled data instances, each unlabeled data instance including values for a plurality of features;

for each feature in the plurality of features, training, by the computer system, a supervised machine learning (ML) model on a labeled dataset derived from the unlabeled dataset, wherein the labeled dataset comprises a plurality of labeled data instances, and wherein each labeled data instance includes:

a label corresponding to a value for the feature in an unlabeled data instance of the unlabeled dataset; and

values for other features in the unlabeled data instance; and

for each pair of first and second features in the plurality of features, computing, by the computer system, an inter-feature influence score using the trained supervised ML model for the second feature, the inter-feature influence score indicating how useful the first feature is in predicting the second feature.

2 . The method of claim 1 wherein computing the inter-feature influence score comprises:

determining a first accuracy score for the trained supervised ML model for the second feature; and

randomly re-shuffling values of the first feature in the labeled dataset for the second feature, resulting in a new labeled dataset for the second feature.

3 . The method of claim 2 wherein computing the inter-feature influence score further comprises:

re-training the supervised ML model for the second feature on the new labeled dataset;

determining a second accuracy score for the re-trained supervised ML model; and

computing the inter-feature influence score based on the first and second accuracy scores.

4 . The method of claim 3 wherein the inter-feature influence score is proportional to a degree of divergence between the first and second accuracy scores.

5 . The method of claim 1 further comprising:

performing dimensionality reduction on the unlabeled dataset based on the inter-feature influence scores.

6 . The method of claim 5 wherein performing dimensionality reduction on the unlabeled dataset comprises:

building a graph comprising a plurality of vertices and a plurality of edges, each vertex corresponding to a feature in the plurality of features, each edge between a pair of vertices being weighted with the inter-feature influence score computed for the pair of features corresponding to the pair of vertices; and

removing all edges in the graph whose weights are below a predefined threshold.

7 . The method of claim 6 wherein performing dimensionality reduction on the unlabeled dataset further comprises, subsequently to the removing:

selecting one or more vertices in the graph from each subset of vertices that remain strongly connected; and

outputting a new unlabeled dataset that includes the plurality of unlabeled data instances but excludes features corresponding to unselected vertices in the graph.

8 . A non-transitory computer readable storage medium having stored thereon program code executable by a computer system, the program code causing the computer system to execute a method comprising:

receiving an unlabeled dataset comprising a plurality of unlabeled data instances, each unlabeled data instance including values for a plurality of features;

for each feature in the plurality of features, training a supervised machine learning (ML) model on a labeled dataset derived from the unlabeled dataset, wherein the labeled dataset comprises a plurality of labeled data instances, and wherein each labeled data instance includes:

a label corresponding to a value for the feature in an unlabeled data instance of the unlabeled dataset; and

values for other features in the unlabeled data instance; and

for each pair of first and second features in the plurality of features, computing an inter-feature influence score using the trained supervised ML model for the second feature, the inter-feature influence score indicating how useful the first feature is in predicting the second feature.

9 . The non-transitory computer readable storage medium of claim 8 wherein computing the inter-feature influence score comprises:

determining a first accuracy score for the trained supervised ML model for the second feature; and

randomly re-shuffling values of the first feature in the labeled dataset for the second feature, resulting in a new labeled dataset for the second feature.

10 . The non-transitory computer readable storage medium of claim 9 wherein computing the inter-feature influence score further comprises:

re-training the supervised ML model for the second feature on the new labeled dataset;

determining a second accuracy score for the re-trained supervised ML model; and

computing the inter-feature influence score based on the first and second accuracy scores.

11 . The non-transitory computer readable storage medium of claim 10 wherein the inter-feature influence score is proportional to a degree of divergence between the first and second accuracy scores.

12 . The non-transitory computer readable storage medium of claim 8 wherein the method further comprises:

performing dimensionality reduction on the unlabeled dataset based on the inter-feature influence scores.

13 . The non-transitory computer readable storage medium of claim 12 wherein performing dimensionality reduction on the unlabeled dataset comprises:

building a graph comprising a plurality of vertices and a plurality of edges, each vertex corresponding to a feature in the plurality of features, each edge between a pair of vertices being weighted with the inter-feature influence score computed for the pair of features corresponding to the pair of vertices; and

removing all edges in the graph whose weights are below a predefined threshold.

14 . The non-transitory computer readable storage medium of claim 13 wherein performing dimensionality reduction on the unlabeled dataset further comprises, subsequently to the removing:

selecting one or more vertices in the graph from each subset of vertices that remain strongly connected; and

outputting a new unlabeled dataset that includes the plurality of unlabeled data instances but excludes features corresponding to unselected vertices in the graph.

15 . A computer system comprising:

a processor; and

a non-transitory computer readable medium having stored thereon program code that, when executed, causes the processor to:

receive an unlabeled dataset comprising a plurality of unlabeled data instances, each unlabeled data instance including values for a plurality of features;

for each feature in the plurality of features, train a supervised machine learning (ML) model on a labeled dataset derived from the unlabeled dataset, wherein the labeled dataset comprises a plurality of labeled data instances, and wherein each labeled data instance includes:

a label corresponding to a value for the feature in an unlabeled data instance of the unlabeled dataset; and

values for other features in the unlabeled data instance; and

for each pair of first and second features in the plurality of features, compute an inter-feature influence score using the trained supervised ML model for the second feature, the inter-feature influence score indicating how useful the first feature is in predicting the second feature.

16 . The computer system of claim 15 wherein the program code that causes the processor to compute the inter-feature influence score comprises program code that causes the processor to:

determine a first accuracy score for the trained supervised ML model for the second feature; and

randomly re-shuffle values of the first feature in the labeled dataset for the second feature, resulting in a new labeled dataset for the second feature.

17 . The computer system of claim 16 wherein the program code that causes the processor to compute the inter-feature influence score further comprises program code that causes the processor to:

re-train the supervised ML model for the second feature on the new labeled dataset;

determine a second accuracy score for the re-trained supervised ML model; and

compute the inter-feature influence score based on the first and second accuracy scores.

18 . The computer system of claim 17 wherein the inter-feature influence score is proportional to a degree of divergence between the first and second accuracy scores.

19 . The computer system of claim 15 wherein the program code further causes the processor to:

perform dimensionality reduction on the unlabeled dataset based on the inter-feature influence scores.

20 . The computer system of claim 19 wherein the program code that causes the processor to perform dimensionality reduction on the unlabeled dataset comprises program code that causes the processor to:

build a graph comprising a plurality of vertices and a plurality of edges, each vertex corresponding to a feature in the plurality of features, each edge between a pair of vertices being weighted with the inter-feature influence score computed for the pair of features corresponding to the pair of vertices; and

remove all edges in the graph whose weights are below a predefined threshold.

21 . The computer system of claim 20 wherein the program code that causes the processor to perform dimensionality reduction on the unlabeled dataset further comprises program code that causes the processor to, subsequently to the removing:

select one or more vertices in the graph from each subset of vertices that remain strongly connected; and

output a new unlabeled dataset that includes the plurality of unlabeled data instances but excludes features corresponding to unselected vertices in the graph.

Assignments (2)
CHANGE OF NAME Recorded Feb 27, 2024
From: VMWARE, INC.
To: VMWARE LLC
Reel/Frame 066692/0103 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2020
From: BEN-ITZHAK, YANIV; VARGAFTIK, SHAY
To: VMWARE, INC.
Reel/Frame 054582/0539 →