IP Library › Granted Patent US 11,604,776
Granted Patent B2
US 11,604,776 · App. 16/839,200 · Granted Mar 14, 2023

Multi-value primary keys for plurality of unique identifiers of entities

Inventors: Andrzej Laskawiec (Cracow, PL); Monika Piatek (Cracow, PL); Lukasz Stanislaw Studzienny (Cracow, PL); Marcin Filip (Cracow, PL); Marcin Luczynski (Cracow, PL); Michal Bodziony (Tęgoborze, PL); Tomasz Zatorski (Cracow, PL)
Assignee: International Business Machines Corporation
G06F16/215G06F16/242G06F16/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,604,776
App. No.
16/839,200
Granted
Mar 14, 2023
Kind
B2
Abstract

A computer-implemented method for unambiguously identifying entities in a database system may be provided. The method comprises storing data items as records with different attributes in a table of a database, storing naming rules for selected combinations of the attributes of the data items, and prioritizing the naming rules. The method also comprises determining a hash value for each of the selected combinations of the attributes of the data items, and identifying duplicate data items using the determined hash values and the prioritized naming rules.

Claims (46)

1. A computer-implemented method for unambiguously identifying entities in a database system, the method comprising:

storing data items in a table of a database, wherein the data items are stored as records comprising a plurality of attributes;

storing naming rules for selected combinations of the attributes of the data items, wherein the naming rules are based on the plurality of attributes including at least, in combination, a customer name and the customer's related employer identification number (EIN);

prioritizing the naming rules by defining a sequence of applying the naming rules or defining an order of the naming rules, depending on an importance of the naming rules for entity detection, wherein the combination of the customer name and the customer's related EIN have a comparably high priority as compared to the customer name alone;

determining a hash value for each of the selected combinations of the attributes of the data items,

identifying duplicate data items using the determined hash values and the prioritized naming rules; and

merging the identified duplicate data items into a merged data item, with the naming rules with higher priority defining where the attributes for the merged data item come from.

2. The method according to claim 1 , wherein the database system is a relational database system, and wherein an entry into the database system automatically triggers a database engine to create the naming rules associated with the entry.

3. The method according to claim 1 , wherein the database system is a configuration management database that underlies a specific internal organization and is used to manage a plurality of technical devices and applications in a plurality of data centers.

4. The method according to claim 1 , further comprising merging the identified duplicate data items by maintaining the determined hash values as a multi-valued key for the merged data item.

5. The method according to claim 4 , further comprising merging other data items that are in composite relationship with the identified data items.

6. The method according to claim 4 , further comprising maintaining a pointer to a same row identifier of one of the merged data items for the determined hash values.

7. The method according to claim 1 , further comprising:

maintaining an index of the table; and

maintaining a pointer in a search tree related to the index, such that the pointer points to the same record identifiers of a combined data item, wherein the combined data item is determined based on splitting a superior alias into two or more individual aliases and each of the two or more individual aliases is used as a single value in the index.

8. The method according to claim 1 , further comprising:

using a create SQL statement adapted for a creating of the naming rule and its related priority.

9. The method according to claim 1 , further comprising:

using a multi-value primary key for sorting records in the table of the database, wherein the multi-value primary key is a unique row identifier.

10. The method according to claim 1 , wherein a multi-value primary key is used for clustering cluster data on multi-node database engines.

11. The method according to claim 9 , wherein a multi-value primary key is comparable to a single value column data item.

12. The method according to claim 1 , further comprising

collecting statistical database data for data blocks for single-valued primary keys and multi-valued primary keys.

13. A computer system for unambiguously identifying entities in the database system, the computer system comprising:

one or more computer processors, one or more computer-readable storage media, and program instructions stored on the one or more of the computer-readable storage media for execution by at least one of the one or more processors capable of performing a method, the method comprising:

storing data items in a table of a database, wherein the data items are stored as records comprising a plurality of attributes;

storing naming rules for selected combinations of the attributes of the data items, wherein the naming rules are based on the plurality of attributes including at least, in combination, a customer name and the customer's related employer identification number (EIN);

prioritizing the naming rules by defining a sequence of applying the naming rules or defining an order of the naming rules, depending on an importance of the naming rules for entity detection, wherein the combination of the customer name and the customer's related EIN have a comparably high priority as compared to the customer name alone;

determining a hash value for each of the selected combinations of the attributes of the data items,

identifying duplicate data items using the determined hash values and the prioritized naming rules; and

merging the identified duplicate data items into a merged data item, with the naming rules with higher priority defining where the attributes for the merged data item come from.

14. The computer system according to claim 13 , wherein the database system is a relational database system.

15. The computer system according to claim 13 , wherein the database system is a configuration management database.

16. The computer system according to claim 13 , further comprising merging the identified duplicate data items by maintaining the determined hash values as a multi-valued key for a merged data item.

17. The computer system according to claim 16 , further comprising merging other data items that are in composite relationship with the identified data items.

18. The computer system according to claim 16 , further comprising maintaining a pointer to a same row identifier of one of the merged data items for the determined hash values.

19. The computer system according to claim 13 , further comprising:

maintaining an index of the table; and

maintaining a pointer in a search tree related to the index, such that the pointer points to the same record identifiers of a combined data item, wherein the combined data item is determined based on splitting a superior alias into two or more individual aliases and each of the two or more individual aliases is used as a single value in the index.

20. A computer program product for unambiguously identifying entities in a database system, the computer program product comprising:

one or more non-transitory computer-readable storage media and program instructions stored on the one or more non-transitory computer-readable storage media capable of performing a method, the method comprising:

storing naming rules for selected combinations of the attributes of the data items, wherein the naming rules are based on the plurality of attributes including at least, in combination, a customer name and the customer's related employer identification number (EIN);

prioritizing the naming rules by defining a sequence of applying the naming rules or defining an order of the naming rules, depending on an importance of the naming rules for entity detection, wherein the combination of the customer name and the customer's related EIN have a comparably high priority as compared to the customer name alone;

determining a hash value for each of the selected combinations of the attributes of the data items,

identifying duplicate data items using the determined hash values and the prioritized naming rules; and

merging the identified duplicate data items into a merged data item, with the naming rules with higher priority defining where the attributes for the merged data item come from.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2020
From: LASKAWIEC, ANDRZEJ; PIATEK, MONIKA; STUDZIENNY, LUKASZ STANISLAW; FILIP, MARCIN; LUCZYNSKI, MARCIN; BODZIONY, MICHAL; ZATORSKI, TOMASZ
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052303/0399 →
Continuity (1)
Related Publication 20210311917A1 · Oct 7, 2021