IP Library Granted Patent US 11,550,834
Granted Patent B1
US 11,550,834 · App. 15/497,484 · Granted Jan 10, 2023

Automated assignment of data set value via semantic matching

Inventors: Stephen Todd (Shrewsbury, MA); David Stephen Reiner (Lexington, MA); Nihar Nanda (Acton, MA)
Assignee: EMC IP Holding Company LLC
G06F16/3344G06F16/93G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,550,834
App. No.
15/497,484
Granted
Jan 10, 2023
Kind
B1
Abstract

An apparatus comprises a processing platform implementing a data set discovery engine and a data set valuation engine. The data set discovery engine is configured to generate data set similarity measures each relating a corresponding one of a plurality of data sets to one or more other ones of the plurality of data sets. The data set valuation engine is coupled to the data set discovery engine and configured to generate valuation measures for respective ones of at least a subset of the plurality of data sets based at least in part on respective ones of the data set similarity measures generated by the data set discovery engine. For example, the data set valuation engine may generate the valuation measure for a given data set as a function of valuation measures previously generated for respective other data sets determined to exhibit at least a threshold similarity to the given data set.

Claims (61)

1. An apparatus comprising:

a processing platform implementing a data set discovery engine and a data set valuation engine;

the data set discovery engine being configured to generate data set similarity measures each relating a corresponding one of a plurality of data sets to one or more other ones of the plurality of data sets; and

the data set valuation engine being coupled to the data set discovery engine and configured to generate valuation measures for respective ones of at least a subset of the plurality of data sets based at least in part on respective ones of the data set similarity measures generated by the data set discovery engine;

the data set similarity measures comprising one or more distance-based similarity measures each measuring a distance between a given one of the plurality of data sets and another one of the plurality of data sets at least in part as a function of a first plurality of data items extracted from the given data set and a second plurality of data items extracted from the other one of the plurality of data sets;

wherein the data set discovery engine in generating the data set similarity measures is configured to semantically examine the given data set using one or more semantic considerations derived from a semantic hierarchy of data sets and relationships to identify similar data sets;

the valuation measures generated for respective ones of the data sets comprising respective quantitative values each representing a quantitative measure of a value of its corresponding data set;

wherein the data set valuation engine in generating the valuation measures is configured:

to generate, for the given data set, a valuation index, the valuation index being generated based at least in part on the given data set and domain-aware tokens associated with the given data set, each of the domain-aware tokens being generated for a corresponding context of interest and comprising information relating to one or more other ones of the data sets that are determined via the semantic examination to be similar to the given data set; and

to determine valuation measures for the given data set by evaluating the valuation index using one or more changing context factors;

wherein the data set valuation engine is configured to assign the valuation measures to respective ones of the data sets by incorporating the valuation measures in metadata of the respective data sets;

wherein the processing platform comprises one or more processing devices each comprising a processor coupled to a memory;

wherein the processing platform further implements at least a portion of a data lake providing one or more of the plurality of data sets for processing by the data set discovery engine and the data set valuation engine; and

wherein at least one of the data sets processed by the data set discovery engine and the data set valuation engine comprises an external data set not yet imported into the data lake and further wherein import of the at least one data set into the data lake is controlled based at least in part on its associated valuation measure.

2. The apparatus of claim 1 wherein the data set valuation engine is configured to generate a valuation measure for the given data set responsive to a received data set valuation request identifying the given data set and values for the one or more changing context factors.

3. The apparatus of claim 2 wherein the data set valuation request is received via a valuation services application programming interface of the data set valuation engine.

4. The apparatus of claim 1 wherein the data set similarity measure generated by the data set discovery engine for at least one of the data sets comprises at least one of distance between the at least one data set and each of one or more other ones of the data sets and a relationship among the at least one data set and one or more other ones of the data sets.

5. The apparatus of claim 1 wherein the data set valuation engine is configured to combine previously-determined valuation measures of the data sets that are determined via the semantic examination to be similar to the given data set by aggregating the previously-determined valuation measures using associated weights.

6. The apparatus of claim 1 wherein at least portions of the semantic hierarchy are adjusted over time through machine learning based at least in part on user interaction with particular data sets and their valuation measures.

7. The apparatus of claim 1 wherein the data set valuation engine is configured to generate the valuation measure for at least one data set as a function of valuation measures previously generated for respective other ones of the data sets determined to exhibit at least a threshold similarity to the at least one data set based at least in part on corresponding ones of the similarity measures generated by the data set discovery engine.

8. The apparatus of claim 1 wherein the data set discovery engine is configured to utilize valuation measures previously assigned to respective ones of the data sets by the data set valuation engine to generate one or more of the data set similarity measures.

9. The apparatus of claim 1 wherein the data set discovery engine is configured:

to generate similarity indexes for the plurality of data sets; and

to obtain a suitability template for a query and to execute the query against one or more of the similarity indexes based at least in part on the suitability template;

wherein the suitability template characterizes suitability for at least one of a particular purpose, a particular goal and a particular role in a process;

wherein the suitability template is associated with at least one target data set and further wherein the data set discovery engine is configured to generate similarity indexes for a plurality of target data sets each associated with one or more suitability templates; and

wherein the suitability template is characterized at least in part by valuation measures of respective data sets.

10. The apparatus of claim 1 wherein the data set valuation engine is configured to adjust valuation measures of respective similar data sets identified by the data set discovery engine.

11. A method comprising:

generating data set similarity measures each relating a corresponding one of a plurality of data sets to one or more other ones of the plurality of data sets; and

generating valuation measures for respective ones of at least a subset of the plurality of data sets based at least in part on respective ones of the data set similarity measures;

the data set similarity measures comprising one or more distance-based similarity measures each measuring a distance between a given one of the plurality of data sets and another one of the plurality of data sets at least in part as a function of a first plurality of data items extracted from the given one of the plurality of data sets and a second plurality of data items extracted from the other one of the plurality of data sets;

the valuation measures generated for respective ones of the data sets comprising respective quantitative values each representing a quantitative measure of a value of its corresponding data set;

wherein the valuation measures are assigned to respective ones of the data sets by incorporating the valuation measures in metadata of the respective data sets;

wherein the generating steps further comprise:

semantically examining the given data set using one or more semantic considerations derived from a semantic hierarchy of data sets and relationships to identify similar data sets;

generating, for the given data set, a valuation index, the valuation index being generated based at least in part on the given data set and domain-aware tokens associated with the given data set, each of the domain-aware tokens being generated for a corresponding context of interest and comprising information relating to one or more other ones of the data sets that are determined via the semantic examination to be similar to the given data set; and

determining valuation measures for the given data set by evaluating the valuation index using one or more changing context factors; and

wherein the method is performed by a processing platform comprising one or more processing devices;

wherein the processing platform implements at least a portion of a data lake providing one or more of the plurality of data sets; and

wherein at least one of the data sets comprises an external data set not yet imported into the data lake and further wherein import of the at least one data set into the data lake is controlled based at least in part on its associated valuation measure.

12. The method of claim 11 wherein generating data set similarity measures comprises utilizing valuation measures previously assigned to respective ones of the data sets to generate one or more of the data set similarity measures.

13. The method of claim 11 wherein generating valuation measures comprises combining previously-determined valuation measures of data sets that are determined via the semantic examination to be similar to the given data set by aggregating the valuation measures using associated weights.

14. The method of claim 11 wherein at least portions of the semantic hierarchy are adjusted over time through machine learning based at least in part on user interaction with particular data sets and their valuation measures.

15. The method of claim 11 wherein a given valuation measure for the given data set is generated responsive to a received data set valuation request identifying the given data set and values for the one or more changing context factors.

16. A computer program product comprising a non-transitory processor-readable storage medium having one or more software programs embodied therein, wherein the one or more software programs when executed by at least one processing device of a processing platform cause the processing device:

to generate data set similarity measures each relating a corresponding one of a plurality of data sets to one or more other ones of the plurality of data sets; and

to generate valuation measures for respective ones of at least a subset of the plurality of data sets based at least in part on respective ones of the data set similarity measures;

the data set similarity measures comprising one or more distance-based similarity measures each measuring a distance between a given one of the plurality of data sets and another one of the plurality of data sets at least in part as a function of a first plurality of data items extracted from the given one of the plurality of data sets and a second plurality of data items extracted from the other one of the plurality of data sets;

the valuation measures generated for respective ones of the data sets comprising respective quantitative values each representing a quantitative measure of a value of its corresponding data set;

wherein the valuation measures are assigned to respective ones of the data sets by incorporating the valuation measures in metadata of the respective data sets;

wherein the one or more software programs when executed by at least one processing device of the processing platform further cause the processing device:

in generating the data set similarity measures, to semantically examine the given data set using one or more semantic considerations derived from a semantic hierarchy of data sets and relationships to identify similar data sets; and

in generating the valuation measures, to generate, for the given data set, a valuation index, the valuation index being generated based at least in part on the given data set and domain-aware tokens associated with the given data set, each of the domain-aware tokens being generated for a corresponding context of interest and comprising information relating to one or more other ones of the data sets that are determined via the semantic examination to be similar to the given data set; and

in generating the valuation measures, to determine valuation measures for the given data set by evaluating the valuation index using one or more changing context factors;

wherein the processing platform implements at least a portion of a data lake providing one or more of the plurality of data sets; and

wherein at least one of the data sets comprises an external data set not yet imported into the data lake and further wherein import of the at least one data set into the data lake is controlled based at least in part on its associated valuation measure.

17. The computer program product of claim 16 wherein generating data set similarity measures comprises utilizing valuation measures previously assigned to respective ones of the data sets to generate one or more of the data set similarity measures.

18. The computer program product of claim 16 wherein generating valuation measures comprises combining previously-determined valuation measures of data sets that are determined via the semantic examination to be similar to the given data set by aggregating the valuation measures using associated weights.

19. The computer program product of claim 16 wherein at least portions of the semantic hierarchy are adjusted over time through machine learning based at least in part on user interaction with particular data sets and their valuation measures.

20. The computer program product of claim 16 wherein a given valuation measure for the given data set is generated responsive to a received data set valuation request identifying the given data set and values for the one or more changing context factors.

Assignments (8)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (042769/0001) Recorded Apr 26, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MOZY, INC.); DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO WYSE TECHNOLOGY L.L.C.)
Reel/Frame 059803/0802 →
RELEASE OF SECURITY INTEREST AT REEL 042768 FRAME 0585 Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; MOZY, INC.; WYSE TECHNOLOGY L.L.C.
Reel/Frame 058297/0536 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2017
From: TODD, STEPHEN; REINER, DAVID STEPHEN; NANDA, NIHAR
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 043665/0512 →
PATENT SECURITY INTEREST (CREDIT) Recorded Jun 12, 2017
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; MOZY, INC.; WYSE TECHNOLOGY L.L.C.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 042768/0585 →
PATENT SECURITY INTEREST (NOTES) Recorded Jun 12, 2017
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; MOZY, INC.; WYSE TECHNOLOGY L.L.C.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 042769/0001 →