IP Library › Granted Patent US 12,361,001
Granted Patent B2
US 12,361,001 · App. 18/617,305 · Granted Jul 15, 2025

Accessing siloed data across disparate locations via a unified metadata graph systems and methods

Inventors: Linfeng Yu (New York, NY); Vaibhav Kumar (New York, NY); Ashutosh Pandey (New York, NY)
Assignee: CITIBANK, N.A.
G06F16/24545G06F16/26G06F16/9024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,001
App. No.
18/617,305
Granted
Jul 15, 2025
Kind
B2
Abstract

Systems and methods for reducing usage of computational resources when accessing siloed data across disparate locations via a unified metadata graph are disclosed. The system receives a user-specified query indicating a request to access a set of data objects. The system then performs natural language processing on the user-specified query to determine a set of phrases corresponding to the user-specified query. The system then accesses a metadata graph to determine a node corresponding to the set of phrases. Using a location identifier corresponding to the determined node, the system determines a data silo storing at least one data object of the set of data objects. The system then generates for display, on a graphical user interface, a visual representation of the at least one data object.

Claims (94)

1. A system for reducing usage of computational resources when accessing siloed data across disparate locations via a unified metadata graph, the system comprising:

at least one hardware processor; and

at least one non-transitory memory storing instructions, which, when executed by the at least one hardware processor, cause the system to:

identifying a set of keywords associated with a request to access a set of data objects;

performing natural language processing on the set of keywords to determine a set of semantically similar phrases corresponding to each keyword of the set of keywords;

accessing a metadata graph to determine a node corresponding to the set of semantically similar phrases, wherein the metadata graph comprises (i) a set of nodes indicating (a) metadata of internal data objects stored in data silos and (b) location identifiers of the data silos, and (ii) edges indicating a data lineage between a first node and a second node of the set of nodes, wherein the metadata graph is generated using a metadata data structure that is based on file-level and container-level metadata identifiers;

determining a data silo storing at least one data object of the set of data objects using the location identifier corresponding to the determined node to obtain the at least one data object of the set of data objects via the data silo; and

generating, for display, on a graphical user interface (GUI), a visual representation of the at least one data object, wherein the visual representation of the at least one data object comprises lineage information of the at least one data object.

2. The system of claim 1 , wherein the metadata graph is generated by:

retrieving (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers from a second set of data silos, wherein each file-level metadata identifier of the set of file-level metadata identifiers indicates metadata of a given data object stored within a respective data silo, and wherein each container-level metadata identifier of the set of container-level metadata identifiers indicates metadata of the respective data silo of the second set of data silos;

generating a set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifiers, respectively;

generating the metadata data structure to map each semantically similar metadata identifier of the set of semantically similar metadata identifiers to normalized file-level metadata identifiers and normalized container-level metadata identifiers; and

generating the metadata graph using the generated metadata data structure.

3. The system of claim 1 , further comprising the instructions to:

receiving, via a second the GUI, a second user-specified query indicating a request to generate an intended result;

providing the second user-specified query to an artificial intelligence model to generate a recommendation, wherein the recommendation comprises (i) a second artificial intelligence model to be used to generate the intended result and (ii) a second set of data objects to be used when training the second artificial intelligence model;

in response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model and (ii) obtaining the second set of data objects using the metadata graph;

training the second artificial intelligence model using the set of data objects; and

applying the second artificial intelligence model to generate the intended result.

4. The system of claim 3 , further comprising the instructions to:

accessing a governance database to obtain a set of policies indicating usage criteria corresponding to the second set of data objects;

determining whether the second set of data objects are approved to be used to train the second artificial intelligence model using the set of policies indicating usage criteria corresponding to the set of second data objects;

determining whether an output of the second artificial intelligence model is approved to be provided to one or more computing systems using a second set of policies indicating usage criteria corresponding to artificial intelligence model predictions; and

in response to (i) the second set of data objects being approved to be used to train the second artificial intelligence model and (ii) the output of the second artificial intelligence model is approved to be provided to one or more computing systems, applying the second artificial intelligence model to generate the intended result.

5. A method for reducing usage of computational resources when accessing siloed data across disparate locations via a unified metadata graph, the method comprising:

identifying a set of keywords associated with a user-specified query to access a set of data objects;

performing natural language processing on the user-specified query to determine a set of phrases corresponding to the user-specified query;

accessing a metadata graph to determine a node corresponding to the set of phrases, wherein the metadata graph comprises (i) a set of nodes comprising (a) metadata indicating internal data objects stored in data silos and (b) location identifiers of the data silos, and (ii) edges indicating data lineages of the set of nodes, wherein the metadata graph is generated using a metadata data structure that is based on file-level and container-level metadata identifiers;

determining a data silo storing at least one data object of the set of data objects using the location identifier corresponding to the determined node to obtain the at least one data object of the set of data objects via the data silo; and

generating a representation of the at least one data object.

6. The method of claim 5 , wherein the metadata graph is generated by:

retrieving (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers from a second set of data silos, wherein each file-level metadata identifier of the set of file-level metadata identifiers indicates metadata of a given data object stored within a respective data silo, and wherein each container-level metadata identifier of the set of container-level metadata identifiers indicates metadata of the respective data silo of the second set of data silos;

generating a set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifiers, respectively;

generating the metadata data structure to map each semantically similar metadata identifier of the set of semantically similar metadata identifiers to normalized file-level metadata identifiers and normalized container-level metadata identifiers; and

generating the metadata graph using the generated metadata data structure.

7. The method of claim 5 , further comprising:

receiving, via a second GUI, a second user-specified query indicating a request to generate an intended result;

providing the second user-specified query to an artificial intelligence model to generate a recommendation, wherein the recommendation comprises (i) a second artificial intelligence model to be used to generate the intended result and (ii) a second set of data objects to be used when training the second artificial intelligence model;

in response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model and (ii) obtaining the second set of data objects using the metadata graph;

training the second artificial intelligence model using the set of data objects; and

applying the second artificial intelligence model to generate the intended result.

8. The method of claim 7 , further comprising:

accessing a governance database to obtain a set of policies indicating usage criteria corresponding to the second set of data objects;

determining whether the second set of data objects are approved to be used to train the second artificial intelligence model using the set of policies indicating usage criteria corresponding to the set of second data objects;

determining whether an output of the second artificial intelligence model is approved to be provided to one or more computing systems using a second set of policies indicating usage criteria corresponding to artificial intelligence model predictions; and

in response to (i) the second set of data objects being approved to be used to train the second artificial intelligence model and (ii) the output of the second artificial intelligence model is approved to be provided to the one or more computing systems, applying the second artificial intelligence model to generate the intended result.

9. The method of claim 5 , wherein determining the set of phrases corresponding to the user-specified query further comprises:

parsing the user-specified query for a set of keywords, wherein each keyword of the set of keywords is associated with the set of data objects;

for each keyword of the set of keywords associated with the set of data objects, determining a set of semantically similar phrases corresponding to the respective keyword of the set of keywords; and

determining the set of phrases corresponding to the user-specified query using the set of semantically similar phrases corresponding to each keyword of the set of keywords.

10. The method of claim 9 , wherein determining a semantically similar phrase corresponding to the respective keyword of the set of keywords further comprises:

accessing a database indicating a mapping between first keywords and a set of second keywords; and

in response to accessing the database, determining the set of semantically similar phrases corresponding to the respective keyword using the respective keyword.

11. The method of claim 5 , wherein accessing the metadata graph further comprises:

traversing each node of the set of nodes to identify a metadata identifier matching at least one phrase of the set of phrases; and

in response to determining that the metadata identifier matches the at least one phrase of the set of phrases, determining the node corresponding to the set of phrases.

12. The method of claim 5 , wherein accessing the metadata graph further comprises:

traversing each node of the set of nodes to identify a metadata identifier matching at least one phrase of the set of phrases;

in response to determining that the metadata identifier matches at least one phrase of the set of phrases, determining a first node corresponding to the set of phrases;

in response to determining the first node corresponds to the set of phrases, performing a second traversal of the nodes of the set of nodes using an edge indicating a first data lineage of the first node, wherein the first data lineage of the first node indicates a second node that comprises information that is a source of information associated with the first node;

determining a second data silo storing a second data object of the set of data objects using the location identifier corresponding to the second node to obtain the second data object of the set of data objects via the second data silo; and

generating a second representation of the second data object.

13. The method of claim 5 , wherein the representation of the at least one data object comprises lineage information of the at least one data object.

14. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause operations comprising:

performing natural language processing on a user-specified query to determine a set of phrases corresponding the user-specified query;

accessing a metadata graph to determine a node corresponding to the set of phrases, wherein the metadata graph comprises (i) a set of nodes comprising (a) metadata indicating internal data objects stored in data silos and (b) location identifiers of the data silos, and (ii) edges indicating data lineages of the set of nodes, wherein the metadata graph is generated using a metadata data structure that is based on file-level and container-level metadata identifiers;

determining a data silo storing at least one data object of the set of data objects using the location identifier corresponding to the determined node to obtain the at least one data object of the set of data objects via the data silo; and

generating a representation of the at least one data object.

15. The media of claim 14 , wherein the metadata graph is generated by:

retrieving (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers from a second set of data silos, wherein each file-level metadata identifier of the set of file-level metadata identifiers indicates metadata of a given data object stored within a respective data silo, and wherein each container-level metadata identifier of the set of container-level metadata identifiers indicates metadata of the respective data silo of the second set of data silos;

generating a set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifiers, respectively;

generating the metadata data structure to map each semantically similar metadata identifier of the set of semantically similar metadata identifiers to normalized file-level metadata identifiers and normalized container-level metadata identifiers; and

generating the metadata graph using the generated metadata data structure.

16. The media of claim 14 , wherein the instructions, when executed by the one or more processors, further cause operations comprising:

receiving, via a second GUI, a second user-specified query indicating a request to generate an intended result;

providing the second user-specified query to an artificial intelligence model to generate a recommendation, wherein the recommendation comprises (i) a second artificial intelligence model to be used to generate the intended result and (ii) a second set of data objects to be used when training the second artificial intelligence model;

in response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model and (ii) obtaining the second set of data objects using the metadata graph;

training the second artificial intelligence model using the set of data objects; and

applying the second artificial intelligence model to generate the intended result.

17. The media of claim 16 , wherein the instructions, when executed by the one or more processors, further cause operations comprising:

accessing a governance database to obtain a set of policies indicating usage criteria corresponding to the second set of data objects;

determining whether the second set of data objects are approved to be used to train the second artificial intelligence model using the set of policies indicating usage criteria corresponding to the set of second data objects;

determining whether an output of the second artificial intelligence model is approved to be provided to one or more computing systems using a second set of policies indicating usage criteria corresponding to artificial intelligence model predictions; and

in response to (i) the second set of data objects being approved to be used to train the second artificial intelligence model and (ii) the output of the second artificial intelligence model is approved to be provided to the one or more computing systems, applying the second artificial intelligence model to generate the intended result.

18. The media of claim 14 , wherein determining the set of phrases corresponding to the user-specified query further comprises:

parsing the user-specified query for a set of keywords, wherein each keyword of the set of keywords is associated with the set of data objects;

for each keyword of the set of keywords associated with the set of data objects, determining a set of semantically similar phrases corresponding to the respective keyword of the set of keywords; and

determining the set of phrases corresponding to the user-specified query using the set of semantically similar phrases corresponding to each keyword of the set of keywords.

19. The media of claim 18 , wherein determining a semantically similar phrase corresponding to the respective keyword of the set of keywords further comprises:

accessing a database indicating a mapping between first keywords and a set of second keywords; and

in response to accessing the database, determining the set of semantically similar phrases corresponding to the respective keyword using the respective keyword.

20. The media of claim 14 , wherein accessing the metadata graph further comprises:

traversing each node of the set of nodes to identify a metadata identifier matching at least one phrase of the set of phrases; and

in response to determining that the metadata identifier matches the at least one phrase of the set of phrases, determining the node corresponding to the set of phrases.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2024
From: YU, LINFENG; KUMAR, VAIBHAV; PANDEY, ASHUTOSH
To: CITIBANK, N.A.
Reel/Frame 067850/0611 →
Continuity (2)
Continuation 18390916 · Dec 20, 2023
Related Publication 20250209074A1 · Jun 26, 2025
References Cited (21)
US 9043372B2 · Makkar et al. · 2015 [cited by applicant]
US 10303688B1 · Sirin et al. · 2019 [cited by applicant]
US 10445170B1 · Subramanian et al. · 2019 [cited by applicant]
US 11138206B2 · Siebeking et al. · 2021 [cited by applicant]
US 11734365B1 · Gottlob et al. · 2023 [cited by applicant]
US 11816154B2 · Ericson · 2023 [cited by applicant]
US 11971891B1 · Yu · 2024 [cited by examiner]
US 12034801B1 · Nair et al. · 2024 [cited by applicant]
US 12045610B1 · Myers et al. · 2024 [cited by applicant]
US 20100198804A1 · Yaskin et al. · 2010 [cited by applicant]
US 20120310975A1 · Oliver · 2012 [cited by examiner]
US 20170091020A1 · Rat et al. · 2017 [cited by applicant]
US 20190235921A1 · Kurian et al. · 2019 [cited by applicant]
US 20200341754A1 · Kunjuramanpillai et al. · 2020 [cited by applicant]
US 20200356725A1 · Okonkwo et al. · 2020 [cited by applicant]
US 20210344745A1 · Mermoud et al. · 2021 [cited by applicant]
US 20220327119A1 · Gasper et al. · 2022 [cited by applicant]
US 20220398498A1 · Vogeti et al. · 2022 [cited by applicant]
US 20230350929A1 · Hasan et al. · 2023 [cited by applicant]
Seabolt, E., et al., “Contextual Intelligence for Unified Data Governance,” In aiDM'IS, Jun. 10, 2018, Houston, TX, USA. ACM, New York, NY, USA, 9 pages. https://doi.org/10.1145/3211954.3211955. [cited by applicant]
Zhihan Lv, Liang Qiao, Sahil Verma, and Kavita. 2021. AI-enabled IoT-Edge Data Analytics for Connected Living. ACM Trans. Internet Technol. 21,4, Article 104,20 pages, <https://doi.org/10.1145/3421510>, (Jun. 2021). [cited by applicant]