IP Library › Granted Patent US 11,989,964
Granted Patent B2
US 11,989,964 · App. 17/524,157 · Granted May 21, 2024

Techniques for graph data structure augmentation

Inventors: Amit Agarwal (Kolkata, IN); Kulbhushan Pachauri (Bangalore, IN); Iman Zadeh (Los Angeles, CA); Jun Qian (Bellevue, WA)
Assignee: Oracle International Corporation
G06V30/41G06N20/00G06V30/18181
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,989,964
App. No.
17/524,157
Filed
Nov 11, 2021
Granted
May 21, 2024
Kind
B2
Art Unit
2683
USPC
382/190
Abstract

A computing device may receive a set of user documents. Data may be extracted from the documents to generate a first graph data structure with one or more initial graphs containing key-value pairs. A model may be trained on the first graph data structure to classify the pairs. Until a set of evaluation metrics for the model exceeds a set of deployment thresholds: generating, a set of evaluation metrics may be generated for the model. The set of evaluation metrics may be compared to the set of deployment thresholds. In response to a determination that the set of evaluation metrics are below the set of deployment thresholds: one or more new graphs may be generated from the one or more initial graphs in the first graph data structure to produce a second graph data structure. The first and second graph can be used to train the model.

Claims (97)

1. A computer-implemented method, comprising:

receiving, at a computing device, a set of user documents;

extracting, by the computing device, data from the set of user documents;

generating, by the computing device, a first graph data structure with one or more initial graphs containing the data extracted from the set of user documents, the data including a set of key-value pairs;

training, by the computing device, a model on the first graph data structure to classify the set of key-value pairs;

until a set of evaluation metrics for the model exceeds a set of deployment thresholds:

generating, by the computing device, the set of evaluation metrics for the model;

comparing, by the computing device, the set of evaluation metrics to the set of deployment thresholds; and

in response to a determination that the set of evaluation metrics are below the set of deployment thresholds:

generating, by the computing device, one or more new graphs from the one or more initial graphs in the first graph data structure to produce a second graph data structure; and

training, by the computing device, the model on the second graph data structure.

2. The method of claim 1 , wherein generating the one or more new graphs to produce the second graph data structure further comprises:

receiving, by the computing device, the one or more initial graphs from the first graph data structure;

deleting, by the computing device, one or more edges and nodes in one or more graphs of the one or more initial graphs to produce one or more new graphs; and

storing, by the computing device, the one or more new graphs in the first graph data structure to produce the second graph data structure.

3. The method of claim 2 , wherein deleting one or more edges and nodes further comprises:

determining, by the computing device, an occurrence metric, the occurrence metric comprising a frequency distribution of a set of node labels;

determining, by the computing device, an importance metric, the importance metric comprising a weight for the set of node labels;

determining, by the computing device, a proximity metric; and

deleting, by the computing device, one or more edges and nodes based at least in part on one or more of the occurrence metric, the importance metric or the proximity metric.

4. The method of claim 1 , further comprising:

until a set of robustness metrics for the model exceeds a set of robustness thresholds;

generating, by the computing device, the set of robustness metrics for the model;

comparing, by the computing device, the set of robustness metrics to the set of robustness thresholds; and

in response to a determination that the set of robustness metrics are below the set of robustness thresholds:

altering, by the computing device, key-value pairs in the second graph data structure to produce a third graph data structure; and

training, by the computing device, the model on at least one of the first graph data structure, the second graph data structure, or the third graph data structure to classify key-value pairs.

5. The method of claim 4 , wherein altering the key-value pairs in the second graph data structure further comprises:

changing, by the computing device, a sequence of words in one or more key-value pairs in the second graph data structure; and

changing, by the computing device, a spelling of one or more words in one or more key-value pairs in the second graph data structure.

6. The method of claim 1 , wherein the user documents include at least one of: drivers licenses, gun licenses, passports, bank cards, employee identification (ID) card, college identification (ID) card, or checks.

7. The method of claim 1 , wherein the data is extracted from the user documents using optical character recognition (OCR).

8. A non-transitory computer-readable storage medium storing a set of instructions, that, when executed by one or more processors of a computer device, cause the one or more processors to perform:

receiving a set of user documents;

extracting data from the set of user documents;

generating a first graph data structure with one or more initial graphs containing the data extracted from the set of user documents, the data including a set of key-value pairs;

training a model on the first graph data structure to classify the set of key-value pairs;

until a set of evaluation metrics for the model exceeds a set of deployment thresholds:

generating the set of evaluation metrics for the model;

comparing the set of evaluation metrics to the set of deployment thresholds; and

in response to a determination that the set of evaluation metrics are below the set of deployment thresholds:

generating one or more new graphs from the one or more initial graphs in the first graph data structure to produce a second graph data structure; and

training the model on the second graph data structure.

9. The computer-readable storage medium of claim 8 , wherein generating the one or more new graphs to produce the second graph data structure further comprises:

receiving the one or more initial graphs from the first graph data structure;

deleting one or more edges and nodes in one or more graphs of the one or more initial graphs to produce one or more new graphs; and

storing the one or more new graphs in the first graph data structure to produce the second graph data structure.

10. The computer-readable storage medium of claim 9 , wherein deleting one or more edges and nodes further comprises:

determining an occurrence metric, the occurrence metric comprising a frequency distribution of a set of node labels;

determining an importance metric, the importance metric comprising a weight for the set of node labels;

determining a proximity metric; and

deleting one or more edges and nodes based at least in part on one or more of the occurrence metric, the importance metric or the proximity metric.

11. The computer-readable storage medium of claim 8 , wherein the set of instructions further comprise instructions for:

until a set of robustness metrics for the model exceeds a set of robustness thresholds;

generating the set of robustness metrics for the model;

comparing the set of robustness metrics to the set of robustness thresholds; and

in response to a determination that the set of robustness metrics are below the set of robustness thresholds:

altering key-value pairs in the first or second graph data structure to produce a third graph data structure; and

training the model on at least one of the first graph data structure, the second graph data structure, or the third graph data structure to classify key-value pairs.

12. The computer-readable storage medium of claim 11 , wherein altering the key-value pairs in either the first or second graph data structure further comprises:

changing a sequence of words in one or more key-value pairs in at least one of the first or second graph data structure; and

changing a spelling of one or more words in one or more key-value pairs in the second graph data structure.

13. The computer-readable storage medium of claim 8 , wherein the user documents include at least one of drivers licenses, gun licenses, passports, bank cards, employee identification (ID) card, college identification (ID) card, or checks.

14. The computer-readable storage medium of claim 8 , wherein the data is extracted from the user documents using optical character recognition (OCR).

15. A system, comprising:

memory storing computer-executable instructions; and

one or more processors configured to access the memory, and execute the computer-executable instructions to at least:

receive a set of user documents;

extract data from the set of user documents;

generate a first graph data structure with one or more initial graphs containing the data extracted from the set of user documents, the data including a set of key-value pairs;

train a model on the first graph data structure to classify the set of key-value pairs;

until a set of evaluation metrics for the model exceeds a set of deployment thresholds:

generate the set of evaluation metrics for the model;

compare the set of evaluation metrics to the set of deployment thresholds; and

in response to a determination that the set of evaluation metrics are below the set of deployment thresholds:

generate one or more new graphs from the one or more initial graphs in the first graph data structure to produce a second graph data structure; and

train the model on the second graph data structure.

16. The system of claim 15 , wherein generating the one or more new graphs to produce the second graph data structure further comprises instructions to:

receive the one or more initial graphs from the first graph data structure;

delete one or more edges and nodes in one or more graphs of the one or more initial graphs to produce one or more new graphs; and

store the one or more new graphs in the first graph data structure to produce the second graph data structure.

17. The system of claim 16 , wherein the instructions to delete one or more edges and nodes further comprises instructions to:

determine an occurrence metric, the occurrence metric comprising a frequency distribution of a set of node labels;

determine an importance metric, the importance metric comprising a weight for the set of node labels;

determine a proximity metric; and

delete one or more edges and nodes based at least in part on one or more of the occurrence metric, the importance metric or the proximity metric.

18. The system of claim 15 , further comprising instructions to:

until a set of robustness metrics for the model exceeds a set of robustness thresholds;

generate, the set of robustness metrics for the model;

compare the set of robustness metrics to the set of robustness thresholds; and

in response to a determination that the set of robustness metrics are below the set of robustness thresholds:

alter key-value pairs in the second graph data structure to produce a third graph data structure; and

train the model on at least one of the first graph data structure, the second graph data structure, or the third graph data structure to classify key-value pairs.

19. The system of claim 18 , wherein altering the key-value pairs in the second graph data structure further comprises instructions to:

change a sequence of words in one or more key-value pairs in the second graph data structure; and

change a spelling of one or more words in one or more key-value pairs in the second graph data structure.

20. The system of claim 15 , wherein the user documents include at least one of: drivers licenses, gun licenses, passports, bank cards, employee identification (ID) card, college identification (ID) card, or checks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2021
From: AGARWAL, AMIT; PACHAURI, KULBHUSHAN; ZADEH, IMAN; QIAN, JUN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 058103/0280 →
Continuity (1)
Related Publication 20230146501A1 · May 11, 2023
Cited By (3)
US 12,602,547 US 12,731,424 US 12,748,874