IP Library Granted Patent US 12,111,870
Granted Patent B2
US 12,111,870 · App. 17/213,946 · Granted Oct 8, 2024

Automatic discovery of related data records

Inventors: Idan Richman Goshen (Beer Sheva, IL); Avitan Gefen (Tel Aviv, IL)
Assignee: EMC IP Holding Company LLC
G06F16/906G06F16/9035G06F16/9038H04L45/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,111,870
App. No.
17/213,946
Granted
Oct 8, 2024
Kind
B2
Abstract

Techniques are provided for automatic discovery of data records. One method comprises obtaining data records each corresponding to a different item and comprising features extracted from a data source, wherein the data records identify related items identified using a collaborative filter that relates items based on user preferences; generating an item network comprising multiple nodes each corresponding to a different item, where two nodes are connected by an edge based on: (i) an item type of the two nodes, (ii) a ratio of numerical values associated with the two nodes, and/or (iii) a pairwise configuration similarity score for the two nodes; clustering the nodes into node clusters based on topological properties of the item network; and identifying items related to a given item that (i) share an edge with the given item and (ii) are in a node cluster comprising a node of the given item.

Claims (38)

1. A method, comprising:

obtaining a plurality of data records, wherein each data record corresponds to a different one of a plurality of items and comprises a plurality of features extracted from at least one data source, wherein at least one data record associated with a first item identifies at least one related item that is related to the first item, wherein the at least one related item is identified in the plurality of data records using a collaborative filter that relates at least some of the items of the plurality of items based at least in part on preferences of a plurality of users, and wherein the collaborative filter identifies, for a given item, one or more additional items obtained or researched by one or more users that also obtained or researched, respectively, the given item;

generating, using the plurality of data records, an item network comprising a plurality of nodes, wherein each node in the item network corresponds to a different one of the plurality of items, wherein two nodes in the item network are selectively connected by an edge in response to an evaluation of: (i) an item type of the items associated with the two nodes, (ii) a ratio of numerical values associated with the two nodes, and (iii) a pairwise configuration similarity score for the two nodes, wherein the pairwise configuration similarity score for the two nodes is based at least in part on a similarity analysis of a textual description of a configuration of each of the items associated with the two nodes, extracted from the at least one data source, for each of the two nodes, wherein the two nodes in the item network are selectively connected by the edge in response to the evaluation determining that: (i) the respective item types of the items associated with the two nodes satisfy one or more similarity criteria, (ii) the ratio of the numerical values associated with the two nodes satisfies a first designated threshold, and (iii) the pairwise configuration similarity score for the two nodes satisfies a second designated threshold, wherein the first designated threshold and the second designated threshold are distinct and wherein the ratio of the numerical values is distinct from the pairwise configuration similarity score;

clustering the plurality of nodes in the item network into a plurality of node clusters based at least in part on an analysis of one or more topological properties of the item network;

identifying one or more items related to a given item by querying the item network to return the one or more identified related items having a corresponding node in the item network that (i) shares an edge with a node in the item network corresponding to the given item and (ii) are in at least one node cluster comprising a node corresponding to the given item; and

initiating an automated processing of at least a given one of the plurality of data records associated with the given item using at least some of the identified one or more items related to the given item;

wherein the method is performed by at least one processing device comprising a processor coupled to a memory.

2. The method of claim 1 , wherein the plurality of items comprises a plurality of products and wherein the features extracted from the at least one data source comprise one or more of a product type, a product name, a product price, a product configuration and a product family.

3. The method of claim 1 , wherein the plurality of items comprises a plurality of products and wherein the plurality of features is extracted from the at least one data source for one or more additional products provided by competitors of a provider of a given product.

4. The method of claim 1 , wherein the plurality of items comprises a plurality of products and wherein the collaborative filter identifies, for a given product, one or more additional products purchased or researched by customers that also purchased or researched, respectively, the given product.

5. The method of claim 1 , wherein the plurality of items comprises a plurality of products and wherein the two nodes in the item network are connected by the edge in response to the two corresponding products having a same product type and having a price ratio that satisfies one or more pricing criteria.

6. The method of claim 1 , further comprising adding one or more edges to the item network using a prediction model trained using one or more features of the item network extracted from the item network, wherein the trained prediction model identifies topological link patterns in the item network to predict at least one missing edge to add to the item network, wherein a weight of the at least one added edge is based at least in part on the pairwise configuration similarity score for the two nodes connected by the at least one added edge.

7. The method of claim 1 , wherein the nodes in a given cluster are more closely related to the nodes in the given cluster than to the nodes in other clusters.

8. The method of claim 1 , wherein the similarity analysis of the textual description of the configuration of each of the items associated with the two nodes comprises one or more of determining a Jaccard similarity and determining a cosine similarity of the configuration of each of the items associated with the two nodes.

9. The method of claim 1 , wherein the plurality of items comprises a plurality of products and wherein the identifying one or more items related to the given item comprises identifying, for a given product, one or more additional products that: (i) are associated with nodes in the item network that share an edge with the node associated with the given product and (ii) are found in the same cluster as the given product.

10. The method of claim 1 , further comprising querying the item network for a particular item of interest to a particular organization, wherein the query returns one or more items that (i) share an edge with a node in the item network corresponding to the particular item, wherein the one or more items that share the edge with the node corresponding to the particular item comprise items competing with the particular item of interest and (ii) are in at least one node cluster comprising a node corresponding to the particular item, wherein the one or more items in the at least one node cluster comprise similar items competing with the particular item of interest.

11. The method of claim 10 , further comprising identifying whether the one or more items returned by the query are provided by one or more of the particular organization and a different organization.

12. The method of claim 1 , wherein the evaluation of the pairwise configuration similarity score for the two nodes is performed in response to the evaluation determining that the respective item types of the items associated with the two nodes satisfy the one or more similarity criteria and the ratio of the numerical values associated with the two nodes satisfies the first designated threshold.

13. An apparatus comprising:

at least one processing device comprising a processor coupled to a memory;

the at least one processing device being configured to implement the following steps:

obtaining a plurality of data records, wherein each data record corresponds to a different one of a plurality of items and comprises a plurality of features extracted from at least one data source, wherein at least one data record associated with a first item identifies at least one related item that is related to the first item, wherein the at least one related item is identified in the plurality of data records using a collaborative filter that relates at least some of the items of the plurality of items based at least in part on preferences of a plurality of users, and wherein the collaborative filter identifies, for a given item, one or more additional items obtained or researched by one or more users that also obtained or researched, respectively, the given item;

generating, using the plurality of data records, an item network comprising a plurality of nodes, wherein each node in the item network corresponds to a different one of the plurality of items, wherein two nodes in the item network are selectively connected by an edge in response to an evaluation of: (i) an item type of the items associated with the two nodes, (ii) a ratio of numerical values associated with the two nodes, and (iii) a pairwise configuration similarity score for the two nodes, wherein the pairwise configuration similarity score for the two nodes is based at least in part on a similarity analysis of a textual description of a configuration of each of the items associated with the two nodes, extracted from the at least one data source, for each of the two nodes, wherein the two nodes in the item network are selectively connected by the edge in response to the evaluation determining that: (i) the respective item types of the items associated with the two nodes satisfy one or more similarity criteria, (ii) the ratio of the numerical values associated with the two nodes satisfies a first designated threshold, and (iii) the pairwise configuration similarity score for the two nodes satisfies a second designated threshold, wherein the first designated threshold and the second designated threshold are distinct and wherein the ratio of the numerical values is distinct from the pairwise configuration similarity score;

clustering the plurality of nodes in the item network into a plurality of node clusters based at least in part on an analysis of one or more topological properties of the item network;

identifying one or more items related to a given item by querying the item network to return the one or more identified related items having a corresponding node in the item network that (i) shares an edge with a node in the item network corresponding to the given item and (ii) are in at least one node cluster comprising a node corresponding to the given item; and

initiating an automated processing of at least a given one of the plurality of data records associated with the given item using at least some of the identified one or more items related to the given item.

14. The apparatus of claim 13 , further comprising adding one or more edges to the item network using a prediction model trained using one or more features of the item network extracted from the item network, wherein the trained prediction model identifies topological link patterns in the item network to predict at least one missing edge to add to the item network, wherein a weight of the at least one added edge is based at least in part on the pairwise configuration similarity score for the two nodes connected by the at least one added edge, wherein the one or more features of the item network extracted from the item network comprise one or more of a joint neighbor feature and a centrality of node feature.

15. The apparatus of claim 13 , wherein the similarity analysis of the textual description of the configuration of each of the items associated with the two nodes comprises one or more of determining a Jaccard similarity and determining a cosine similarity of the configuration of each of the items associated with the two nodes.

16. The apparatus of claim 13 , wherein the plurality of items comprises a plurality of products and wherein the identifying one or more items related to the given item comprises identifying, for a given product, one or more additional products that: (i) are associated with nodes in the item network that share an edge with the node associated with the given product and (ii) are found in the same cluster as the given product.

17. A non-transitory processor-readable storage medium having stored therein program code of one or more software programs, wherein the program code when executed by at least one processing device causes the at least one processing device to perform the following steps:

obtaining a plurality of data records, wherein each data record corresponds to a different one of a plurality of items and comprises a plurality of features extracted from at least one data source, wherein at least one data record associated with a first item identifies at least one related item that is related to the first item, wherein the at least one related item is identified in the plurality of data records using a collaborative filter that relates at least some of the items of the plurality of items based at least in part on preferences of a plurality of users, and wherein the collaborative filter identifies, for a given item, one or more additional items obtained or researched by one or more users that also obtained or researched, respectively, the given item;

generating, using the plurality of data records, an item network comprising a plurality of nodes, wherein each node in the item network corresponds to a different one of the plurality of items, wherein two nodes in the item network are selectively connected by an edge in response to an evaluation of: (i) an item type of the items associated with the two nodes, (ii) a ratio of numerical values associated with the two nodes, and (iii) a pairwise configuration similarity score for the two nodes, wherein the pairwise configuration similarity score for the two nodes is based at least in part on a similarity analysis of a textual description of a configuration of each of the items associated with the two nodes, extracted from the at least one data source, for each of the two nodes, wherein the two nodes in the item network are selectively connected by the edge in response to the evaluation determining that: (i) the respective item types of the items associated with the two nodes satisfy one or more similarity criteria, (ii) the ratio of the numerical values associated with the two nodes satisfies a first designated threshold, and (iii) the pairwise configuration similarity score for the two nodes satisfies a second designated threshold, wherein the first designated threshold and the second designated threshold are distinct and wherein the ratio of the numerical values is distinct from the pairwise configuration similarity score;

clustering the plurality of nodes in the item network into a plurality of node clusters based at least in part on an analysis of one or more topological properties of the item network;

identifying one or more items related to a given item by querying the item network to return the one or more identified related items having a corresponding node in the item network that (i) shares an edge with a node in the item network corresponding to the given item and (ii) are in at least one node cluster comprising a node corresponding to the given item; and

initiating an automated processing of at least a given one of the plurality of data records associated with the given item using at least some of the identified one or more items related to the given item.

18. The non-transitory processor-readable storage medium of claim 17 , wherein the plurality of items comprises a plurality of products and wherein the collaborative filter identifies, for a given product, one or more additional products purchased or researched by customers that also purchased or researched, respectively, the given product.

19. The non-transitory processor-readable storage medium of claim 17 , further comprising adding one or more edges to the item network using a prediction model trained using one or more features of the item network extracted from the item network, wherein the trained prediction model identifies topological link patterns in the item network to predict at least one missing edge to add to the item network, wherein a weight of the at least one added edge is based at least in part on the pairwise configuration similarity score for the two nodes connected by the at least one added edge, wherein the one or more features of the item network extracted from the item network comprise one or more of a joint neighbor feature and a centrality of node feature.

20. The non-transitory processor-readable storage medium of claim 17 , wherein the plurality of items comprises a plurality of products and wherein the identifying one or more items related to the given item comprises identifying, for a given product, one or more additional products that: (i) are associated with nodes in the item network that share an edge with the node associated with the given product and (ii) are found in the same cluster as the given product.

Assignments (10)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (056295/0280) Recorded Jun 10, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
Reel/Frame 062022/0255 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (056295/0124) Recorded Jun 10, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
Reel/Frame 062022/0012 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (056295/0001) Recorded Jun 10, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
Reel/Frame 062021/0844 →
RELEASE OF SECURITY INTEREST Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
Reel/Frame 058297/0332 →
SECURITY INTEREST Recorded May 19, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 056295/0124 →
SECURITY INTEREST Recorded May 19, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 056295/0001 →
SECURITY INTEREST Recorded May 19, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
Reel/Frame 056295/0280 →
CORRECTIVE ASSIGNMENT TO CORRECT THE MISSING PATENTS THAT WERE ON THE ORIGINAL SCHEDULED SUBMITTED BUT NOT ENTERED PREVIOUSLY RECORDED AT REEL: 056250 FRAME: 0541. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 17, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 056311/0781 →
SECURITY AGREEMENT Recorded May 14, 2021
From: DELL PRODUCTS L.P.; EMC IP HOLDING COMPANY LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 056250/0541 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2021
From: GOSHEN, IDAN RICHMAN; GEFEN, AVITAN
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 055734/0729 →
Continuity (1)
Related Publication 20220309100A1 · Sep 29, 2022