IP Library Granted Patent US 11,106,642
Granted Patent B2
US 11,106,642 · App. 16/232,568 · Granted Aug 31, 2021

Cataloging database metadata using a probabilistic signature matching process

Inventors: Tomoya Wada (New York, NY); Winnie Cheng (West New York, NJ); Rohit Mahajan (Iselin, NJ); Alex Mylnikov (Edison, NJ)
Assignee: Io-Tahoe LLC.
G06F16/213G06F16/221G06F16/2456G06F16/25G06F16/908G06F16/90332G06F16/90344G06F40/242G06K9/6215G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,106,642
App. No.
16/232,568
Granted
Aug 31, 2021
Kind
B2
Abstract

A system and computer implemented method for cataloging database metadata using a probabilistic signature matching process are provided. The method includes receiving an input name to be matched to keys in a data corpus; dividing the received input name into a plurality of text segments; identifying a set of matching keys by matching each of the plurality text segments against keys in the data corpus; analyzing the set of matching keys to construct a tag; and cataloging the metadata with the matching key as the construct tag.

Claims (43)

1. A computer implemented method for cataloging database metadata using a probabilistic signature matching process, comprising:

receiving an input name to be matched to keys in a data corpus;

dividing the received input name into a plurality of text segments;

identifying a set of matching keys by matching each of the plurality of text segments against keys in the data corpus, wherein the identifying is based on developing for each text segment at least one of a plurality of n-grams, at least two of the plurality of n-grams having different orders and being developed from letters of the received input name and at least two of the plurality of n-grams having different orders and being developed from letters of a phonetic expression of the received input name based on a pronunciation schema;

analyzing the set of matching keys to construct a tag; and

cataloging the database metadata with the constructed tag.

2. The computer implemented method of claim 1 , wherein dividing the received input name further comprises:

dividing the input name into strings;

determining for each string whether a respective string is found in the data corpus; and

determining each string found in the data corpus as a text segment.

3. The computer implemented method of claim 2 , further comprising:

deploying dynamic programming to determine whether the respective string is found in the data corpus.

4. The computer implemented method of claim 1 , wherein matching of each of the plurality of text segments against keys in the data corpus further comprises:

determining, for each of the plurality of text segments, a probability for a matching key, wherein the determination is based on the contents of the data corpus; and

selecting a one of the plurality of text segments with the highest probability as the matching key.

5. The computer implemented method of claim 4 , wherein the probability is a function of at least one of: a frequency of appearance for each of the plurality of text segments in the data corpus and a count of each of the plurality of text segments in the corpus.

6. The computer implemented method of claim 1 , wherein the data corpus includes pairs of keys and values related to input names and corresponding respective tags.

7. The computer implemented method of claim 6 , wherein the data corpus is trained on a broad range of data sources in a specific data domain.

8. The computer implemented method of claim 1 , wherein the tag is a description of the input name.

9. The computer implemented method of claim 1 , wherein the received input name is a column name of one or more databases across multiple data sources.

10. A non-transitory computer readable medium having stored thereon instructions for causing a processing circuitry to execute the method of claim 1 .

11. A system for cataloging database metadata using a probabilistic signature matching process, comprising:

a processing circuitry; and

a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to:

receive an input name to be matched to keys in a data corpus;

divide the received input name into a plurality of text segments;

identify a set of matching keys by matching each of the plurality text segments against keys in the data corpus, wherein the identifying is based on developing for each text segment at least one of a plurality of n-grams, at least two of the plurality of n-grams having different orders and being developed from letters of the received input name and at least two of the plurality of n-grams having different orders and being developed from letters of a phonetic expression of the received input name based on a pronunciation schema;

analyze the set of matching keys to construct a tag; and

catalog the database metadata with the constructed tag.

12. The system of claim 11 , wherein the system is further configured to:

divide the input name into strings;

determine for each string whether a respective string is found in the data corpus; and

determine each string found in the data corpus as a text segment.

13. The system of claim 12 , wherein the system is further configured to:

deploy dynamic programming to determine whether the respective string is found in the data corpus.

14. The system of claim 11 , wherein the system is further configured to:

determine, for each of the plurality of text segments, a probability for a matching key, wherein the determination is based on the contents of the data corpus; and

select a one of the plurality of text segments with the highest probability as the matching key.

15. The system of claim 14 , wherein the probability is a function of at least one of: a frequency of appearance for each of the plurality of text segments in the data corpus and a count of each of the plurality of text segments in the corpus.

16. The system of claim 11 , wherein the data corpus includes pairs of keys and values related to input names and corresponding respective tags.

17. The system of claim 16 , wherein the data corpus is trained on a broad range of data sources in a specific data domain.

18. The computer implemented method of claim 1 , wherein the tag is a description of the input name.

19. The system of claim 11 , wherein the received input name is a column name of one or more databases across multiple data sources.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2021
From: IO-TAHOE LLC
To: HITACHI VANTARA LLC
Reel/Frame 057323/0314 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 26, 2018
From: WADA, TOMOYA; CHENG, WINNIE; MAHAJAN, ROHIT; MYLNIKOV, ALEX
To: IO-TAHOE LLC.
Reel/Frame 047854/0192 →
Continuity (1)
Related Publication 20200210388A1 · Jul 2, 2020