IP Library Granted Patent US 12,353,834
Granted Patent B2
US 12,353,834 · App. 18/056,615 · Granted Jul 8, 2025

Systems and methods for generalized structured data discovery utilizing contextual metadata disambiguation via machine learning techniques

Inventors: Santosh Chikoti (Monroe Township, NJ); Jeffrey Kessler (Mahopac, NY); Saurabh Gupta (Secaucus, NJ); Deepak Jayadas (Newark, DE)
Assignee: JPMORGAN CHASE BANK, N.A.
G06F40/30G06F16/24573G06F40/205G06F40/253G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,353,834
App. No.
18/056,615
Granted
Jul 8, 2025
Kind
B2
Abstract

Systems and methods for generalized structured data discovery utilizing contextual metadata disambiguation via machine learning are disclosed. A method may include receiving physical application metadata for an attribute, a database object, or a database; receiving reference data including tokens and their associated abbreviations and acronyms; parsing the physical application metadata into application tokens including known application tokens and unknown application tokens; identifying unknown application tokens by comparing the parsed application tokens to a corpus; performing probabilistic parsing on the unknown application tokens using the reference data resulting in token phrases; identifying monosemous tokens in the token phrases using a look-up dictionary; replacing the monosemous tokens with their expansions from the look-up dictionary; and outputting a mapping of the physical application metadata to enhanced physical application metadata that includes an expression for the physical application metadata comprising the expansions for monosemous tokens in a supported language.

Claims (38)

1. A method for generalized structured data discovery utilizing contextual metadata disambiguation via machine learning techniques, comprising:

in an information processing apparatus comprising at least one computer processor:

receiving physical application metadata from one or more data sources for an attribute, a database object, or a database;

receiving reference data comprising a plurality of tokens and their associated abbreviations and acronyms;

parsing the physical application metadata into a plurality of application tokens comprising known application tokens and unknown application tokens;

identifying unknown application tokens by comparing the parsed application tokens to a corpus;

performing probabilistic parsing on the unknown application tokens using the reference data resulting in token phrases;

identifying monosemous tokens in the token phrases using a look-up dictionary;

replacing the monosemous tokens with their expansions from the look-up dictionary; and

outputting a mapping of the physical application metadata to enhanced physical application metadata, wherein the enhanced physical application metadata comprises an expression for the physical application metadata comprising the expansions for monosemous tokens in a supported language.

2. The method of claim 1 , further comprising:

performing neural machine translation on unknown application tokens in an unsupported language to translate the unknown application tokens to the supported language.

3. The method of claim 1 , wherein the metadata is received from a plurality of data sources.

4. The method of claim 1 , wherein the parsing performs at least one of the following: (1) eliminates special characters; and (2) removes stop words.

5. The method of claim 1 , wherein the parsing is based on common delimiters.

6. The method of claim 1 , wherein the corpus comprises a dictionary for the supported language.

7. The method of claim 1 , wherein the corpus comprises the reference data.

8. The method of claim 1 , wherein the reference data comprises common industry and organization terms.

9. A system for generalized structured data discovery utilizing contextual metadata disambiguation via machine learning techniques, comprising:

a plurality of data sources of physical application metadata for an attribute, a database object, or a database;

at least one organizational database comprising reference data comprising a plurality of tokens and their associated abbreviations and acronyms; and

a language processing engine comprising at least one computer processor in communication with the plurality of data sources;

wherein:

the language processing engine receives the physical application metadata from the data sources;

the language processing engine receives the reference data from the organizational database;

the language processing engine parses the physical application metadata into a plurality of application tokens comprising known application tokens and unknown application tokens;

the language processing engine identifies unknown application tokens by comparing the parsed application tokens to a corpus;

the language processing engine performs probabilistic parsing on the unknown application tokens using the reference data resulting in token phrases;

the language processing engine identifies monosemous tokens in the token phrases using a look-up dictionary;

the language processing engine replaces the monosemous tokens with their expansions from the look-up dictionary; and

the language processing engine outputs a mapping of the physical application metadata to enhanced physical application metadata, wherein the enhanced physical application metadata comprises an expression for the physical application metadata comprising the expansions for monosemous tokens in a supported language.

10. The system of claim 9 , wherein the language processing engine further performs neural machine translation on unknown application tokens in an unsupported language to translate the unknown application tokens to the supported language.

11. The system of claim 9 , wherein the metadata is received from a plurality of data sources.

12. The system of claim 9 , wherein the parsing performs at least one of the following: (1) eliminates special characters; and (2) removes stop words.

13. The system of claim 9 , wherein the parsing is based on common delimiters.

14. The system of claim 9 , wherein the corpus comprises a dictionary for the supported language.

15. The system of claim 9 , wherein the corpus comprises the reference data.

16. The system of claim 9 , wherein the reference data comprises common industry and organization terms.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: CHIKOTI, SANTOSH; KESSLER, JEFFREY; GUPTA, SAURABH; JAYADAS, DEEPAK
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064087/0778 →
Continuity (2)
Continuation 17010023 · Sep 2, 2020
Related Publication 20230087421A1 · Mar 23, 2023
References Cited (12)
US 10290003B1 · Hammad et al. · 2019 [cited by applicant]
US 10311402B1 · Mathwig et al. · 2019 [cited by applicant]
US 11449676B2 · Reddi et al. · 2022 [cited by applicant]
US 11449842B2 · Baldet et al. · 2022 [cited by applicant]
US 20020133514A1 · Bates · 2002 [cited by examiner]
US 20020180807A1 · Dubil · 2002 [cited by examiner]
US 20120072204A1 · Nasri · 2012 [cited by examiner]
US 20140032584A1 · Baker · 2014 [cited by examiner]
US 20150186363A1 · Vashishtha · 2015 [cited by examiner]
US 20180082032A1 · Allen · 2018 [cited by examiner]
US 20190043500A1 · Malik · 2019 [cited by examiner]
US 20190236131A1 · Allen · 2019 [cited by examiner]