IP Library Granted Patent US 11,263,262
Granted Patent B2
US 11,263,262 · App. 16/654,661 · Granted Mar 1, 2022

Indexing a dataset based on dataset tags and an ontology

Inventor: Kai-Wen Chen (Arlington, VA)
Assignee: Capital One Services, LLC
G06F16/901G06F16/367G06F16/9038
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,263,262
App. No.
16/654,661
Granted
Mar 1, 2022
Kind
B2
Abstract

Aspects described herein may relate to methods, systems, and apparatuses that process one or more tags associated with a dataset and index the dataset based on the processing of the one or more tags. Processing a tag may include, for example, tokenizing the tag, mapping or expanding abbreviations included within the tag, and otherwise mapping or expanding elements of the tag based on alphanumeric characteristics. Additionally, as part of processing the tag, a number of potential tags may be determined. An ontology may be searched to determine whether any of the potential tags are also found within the ontology. The dataset may be indexed into a searchable index based on any of the potential tags that are found within the ontology.

Claims (53)

1. A method comprising:

determining a tag associated with a dataset;

determining, by one or more computing devices, a plurality of tokenized elements by processing the tag based on one or more alphanumeric characteristics for splitting tag elements;

determining, by the one or more computing devices, a plurality of expanded tokenized elements by processing the plurality of tokenized elements based on one or more abbreviation mappings for abbreviations that include alphabetic or numeric abbreviations;

determining, by the one or more computing devices, a plurality of potential ontology tags by processing the plurality of expanded tokenized elements based on a plurality of tag extraction windows;

determining one or more ontology tags by processing the plurality of potential ontology tags based on an ontology that includes a plurality of standardized dataset tags; and

storing an association between the one or more ontology tags and the dataset.

2. The method of claim 1 , wherein the dataset is formatted in one or more columns, and the tag includes a column name for a column of the one or more columns.

3. The method of claim 1 , wherein the steps of claim 1 are performed based on a request to register the dataset with a metadata repository.

4. The method of claim 1 , wherein the tag includes a first element, wherein the first element includes an abbreviation with alphabetic and numeric characters, and wherein the plurality of expanded tokenized elements includes a second element that is based on applying the one or more abbreviation mappings to the first element.

5. The method of claim 1 , wherein the one or more alphanumeric characteristics include an occurrence of camel case or snake case.

6. The method of claim 1 , wherein the one or more alphanumeric characteristics include an occurrence of a transition between an alphabetic character and a numeric character.

7. The method of claim 1 , wherein the one or more alphanumeric characteristics include an occurrence of one or more of a new line character, a tab character, a comma character, or a space character.

8. The method of claim 1 , wherein the one or more alphanumeric characteristics include occurrences of camel case, a transition between an alphabetic character and a numeric character, a new line character, a tab character, a comma character, and a space character.

9. The method of claim 1 , wherein the plurality of tokenized elements includes a first element, wherein the first element includes an abbreviation with alphabetic or numeric characters, and wherein the plurality of expanded tokenized elements includes a second element that is based on applying the one or more abbreviation mappings to the first element.

10. The method of claim 1 , wherein processing the plurality of expanded tokenized elements based on the plurality of tag extraction windows is performed by sliding each of the plurality of tag extraction windows over the plurality of expanded tokenized elements;

wherein the plurality of tag extraction windows includes a first window that has a size of a single element; and

wherein the plurality of tag extraction windows includes a second window that has a size equal to a number of tokenized elements in the plurality of expanded tokenized elements.

11. The method of claim 1 , wherein processing the plurality of potential ontology tags based on the ontology is performed by matching one or more of the plurality of potential ontology tags with one or more of the plurality of standardized dataset tags.

12. The method of claim 1 , wherein the tag includes alphanumeric elements arranged in an order; and

wherein each the plurality of tokenized elements, the plurality of expanded tokenized elements, and the plurality of potential ontology tags maintains the order.

13. The method of claim 1 , further comprising:

receiving a request to search for a first standardized dataset tag of the plurality of standardized dataset tags;

performing, based on the first standardized dataset tag, a search of a searchable index;

causing, based on the search, display of an indication of the dataset;

receiving a request for the dataset; and

causing, based on the request for the dataset, display of the dataset.

14. One or more non-transitory media storing instructions that, when executed, cause one or more computing devices to perform steps comprising:

determining a tag associated with a dataset;

determining a plurality of tokenized elements by processing the tag based on one or more alphanumeric characteristics for splitting tag elements;

determining a plurality of expanded tokenized elements by processing the plurality of tokenized elements based on one or more abbreviation mappings for abbreviations that include alphabetic or numeric abbreviations;

determining a plurality of potential ontology tags by processing the plurality of expanded tokenized elements based on a plurality of tag extraction windows;

determining one or more ontology tags by processing the plurality of potential ontology tags based on an ontology that includes a plurality of standardized dataset tags; and

storing an association between the one or more ontology tags and the dataset.

15. The one or more non-transitory media of claim 14 , wherein the dataset is formatted in one or more columns, and the tag includes a column name for a column of the one or more columns.

16. The one or more non-transitory media of claim 14 , wherein the one or more alphanumeric characteristics include occurrences of camel case, a transition between an alphabetic character and a numeric character, a new line character, a tab character, a comma character, and a space character.

17. The one or more non-transitory media of claim 14 , wherein the plurality of tokenized elements includes a first element, wherein the first element includes an abbreviation with alphabetic or numeric characters, and wherein the plurality of expanded tokenized elements includes a second element that is based on applying the one or more abbreviation mappings to the first element.

18. The one or more non-transitory media of claim 14 , wherein processing the plurality of expanded tokenized elements based on the plurality of tag extraction windows is performed by sliding each of the plurality of tag extraction windows over the plurality of expanded tokenized elements;

wherein the plurality of tag extraction windows includes a first window that has a size of a single element; and

wherein the plurality of tag extraction windows includes a second window that has a size equal to a number of tokenized elements in the plurality of expanded tokenized elements.

19. The one or more non-transitory media of claim 14 , wherein the tag includes alphanumeric elements arranged in an order; and

wherein each of the plurality of tokenized elements, the plurality of expanded tokenized elements, and the plurality of potential ontology tags maintains the order.

20. A system comprising:

one or more first computing devices configured to operate as part of a metadata repository; and

one or more second computing devices that comprise:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the one or more second computing devices to perform steps comprising:

determining, based on a request to register a dataset with the metadata repository, a tag associated with the dataset;

determining a plurality of tokenized elements by processing the tag based on one or more alphanumeric characteristics for splitting tag elements;

determining a plurality of expanded tokenized elements by processing the plurality of tokenized elements based on one or more abbreviation mappings for abbreviations that include alphabetic or numeric abbreviations;

determining a plurality of potential ontology tags by processing the plurality of expanded tokenized elements based on a plurality of tag extraction windows;

determining one or more ontology tags by processing the plurality of potential ontology tags based on an ontology that includes a plurality of standardized dataset tags; and

storing an association between the one or more ontology tags and the dataset.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2020
From: CHEN, KAI-WEN
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 053096/0780 →
Continuity (2)
Continuation 16457706 · Jun 28, 2019
Related Publication 20200409999A1 · Dec 31, 2020
Cited By (18)
US 12,190,330 US 12,204,564 US 12,216,794 US 12,259,882 US 12,265,896 US 12,277,232 US 12,288,233 US 12,299,065 US 12,353,405 US 12,381,915 US 12,412,140 US 12,536,329 US 12,591,828 US 12,609,938 US 12,641,108 US 12,688,324 US 12,694,044 US 12,718,167