IP Library › Granted Patent US 12,106,595
Granted Patent B2
US 12,106,595 · App. 18/379,091 · Granted Oct 1, 2024

Pseudo labelling for key-value extraction from documents

Inventors: Amit Agarwal (Kolkata, IN); Kulbhushan Pachauri (Bangalore, IN)
Assignee: Oracle International Corporation
G06V30/414G06V30/19147G06V30/19173G06V30/19187
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,106,595
App. No.
18/379,091
Granted
Oct 1, 2024
Kind
B2
Abstract

A computing device may access visually rich documents comprising an image and metadata. A graph, based on the image or metadata, can be generated for a visually rich document. The graph's nodes can correspond to words from the visually rich document. Features for nodes can be determined by the device. The device may generate model labeled graphs by assigning a pseudo-label to nodes using a pretrained model. The device may generate a plurality of graph labeled graphs by assigning a pseudo-label to nodes by matching a first node from a first graph to at least a second node from a second graph. The device may generate a plurality of updated graphs by cross referencing labels from the model labeled graphs and the graph labeled graphs. Until a change in labels is below a threshold, a model can be trained to perform key-value extraction using the updated graphs.

Claims (47)

1. A method, comprising:

accessing, by a computing device, two or more graphs, each graph being generated from a visually rich document and each graph comprising a plurality of nodes connected by a plurality of edges;

generating, by a model pseudo-labeling module of the computing device, a plurality of model labeled graphs by assigning a model pseudo-label to at least a subset of the nodes using a pretrained model;

generating, by a graph pseudo-labeling module of the computing device, a plurality of graph labeled graphs by assigning a graph pseudo-label to at least a subset of the nodes by matching a first node from a first graph to at least a second node from a second graph;

generating, by a filtering module of the computing device, a plurality of updated graphs by updating the nodes based at least in part on cross referencing labels from the model labeled graphs and the graph labeled graphs; and

storing, by the computing device, the plurality of updated graphs.

2. The method of claim 1 , wherein generating the plurality of updated graphs further comprises:

identifying, by the filtering module of the computing device, a model labeled graph and a graph labeled graph that correspond to the same visually rich document;

identifying, by the filtering module of the computing device, an inconsistent node where the model pseudo-label and the graph pseudo-label do not match; and

updating, by the filtering module of the computing device, an inconsistent label for the inconsistent node based at least in part on a model confidence score for the model pseudo-label or a graph confidence score for the graph pseudo-label.

3. The method of claim 1 , wherein at least one graph of the two or more graphs is based at least in part on metadata for the visually rich document, the metadata including at least one of a plurality of words identified with optical character recognition (OCR), a set of user-thresholds, or a plurality of labels.

4. The method of claim 1 , wherein at least one graph of the two or more graphs was generated from a labeled visually rich document.

5. The method of claim 1 , wherein the visually rich document includes at least one of: a drivers license, a medical bill, a gun license, a passport, a bank card, an employee identification (ID) card, a college identification (ID) card, an invoice, a receipt, a business card, a product catalog, a bank form, an investment form, a credit card statement, an account statement, an insurance form, a real estate form, a hospital form, a registration form, a proof of delivery document, a shipment bill, an inquiry form or a check.

6. The method of claim 1 , wherein the plurality of features includes at least one of: structural information, textual information, or visual information.

7. The method of claim 1 , wherein the plurality of graph labeled graphs are generated based at least in part on bipartite graph matching.

8. A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising:

one or more instructions that, when executed by one or more processors of a computing device, cause the computing device to:

access two or more graphs, each graph being generated from a visually rich document and each graph comprising a plurality of nodes connected by a plurality of edges;

generate a plurality of model labeled graphs by assigning a model pseudo-label to at least a subset of the nodes using a pretrained model;

generate a plurality of graph labeled graphs by assigning a graph pseudo-label to at least a subset of the nodes by matching a first node from a first graph to at least a second node from a second graph;

generating a plurality of updated graphs by updating the nodes based at least in part on cross referencing labels from the model labeled graphs and the graph labeled graphs; and

store the plurality of updated graphs.

9. The non-transitory computer-readable medium of claim 8 , wherein the one or more instructions, that cause the computing device to generate the plurality of updated graphs, cause the computing device to:

identify a model labeled graph and a graph labeled graph that correspond to the same visually rich document;

identify an inconsistent node where the model pseudo-label and the graph pseudo-label do not match; and

update an inconsistent label for the inconsistent node based at least in part on a model confidence score for the model pseudo-label or a graph confidence score for the graph pseudo-label.

10. The non-transitory computer-readable medium of claim 8 , wherein at least one graph of the two or more graphs is based at least in part on metadata for the visually rich document, the metadata including at least one of a plurality of words identified with optical character recognition (OCR), a set of user-thresholds, or a plurality of labels.

11. The non-transitory computer-readable medium of claim 8 , wherein at least one graph of the two or more graphs was generated from a labeled visually rich document.

12. The non-transitory computer-readable medium of claim 8 , wherein the visually rich document includes at least one of: a drivers license, a medical bill, a gun license, a passport, a bank card, an employee identification (ID) card, a college identification (ID) card, an invoice, a receipt, a business card, a product catalog, a bank form, an investment form, a credit card statement, an account statement, an insurance form, a real estate form, a hospital form, a registration form, a proof of delivery document, a shipment bill, an inquiry form or a check.

13. The non-transitory computer-readable medium of claim 8 , wherein the plurality of features includes at least one of: structural information, textual information, or visual information.

14. The non-transitory computer-readable medium of claim 8 , wherein the plurality of graph labeled graphs are generated based at least in part on bipartite graph matching.

15. A computing device, comprising:

one or more memories; and

one or more processors, communicatively coupled to the one or more memories, configured to:

access two or more graphs, each graph being generated from a visually rich document and each graph comprising a plurality of nodes connected by a plurality of edges;

generate a plurality of model labeled graphs by assigning a model pseudo-label to at least a subset of the nodes using a pretrained model;

generate a plurality of graph labeled graphs by assigning a graph pseudo-label to at least a subset of the nodes by matching a first node from a first graph to at least a second node from a second graph;

generating a plurality of updated graphs by updating the nodes based at least in part on cross referencing labels from the model labeled graphs and the graph labeled graphs; and

store the plurality of updated graphs.

16. The computing device of claim 15 , wherein the one or more processors, when generating the plurality of updated graphs, are configured to:

identify a model labeled graph and a graph labeled graph that correspond to the same visually rich document;

identify an inconsistent node where the model pseudo-label and the graph pseudo-label do not match; and

update an inconsistent label for the inconsistent node based at least in part on a model confidence score for the model pseudo-label or a graph confidence score for the graph pseudo-label.

17. The computing device of claim 15 , wherein at least one graph of the two or more graphs is based at least in part on metadata for the visually rich document, the metadata including at least one of a plurality of words identified with optical character recognition (OCR), a set of user-thresholds, or a plurality of labels.

18. The computing device of claim 15 , wherein at least one graph of the two or more graphs was generated from a labeled visually rich document.

19. The computing device of claim 15 , wherein the visually rich document includes at least one of: a drivers license, a medical bill, a gun license, a passport, a bank card, an employee identification (ID) card, a college identification (ID) card, an invoice, a receipt, a business card, a product catalog, a bank form, an investment form, a credit card statement, an account statement, an insurance form, a real estate form, a hospital form, a registration form, a proof of delivery document, a shipment bill, an inquiry form or a check.

20. The computing device of claim 15 , wherein the plurality of features includes at least one of: structural information, textual information, or visual information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2023
From: AGARWAL, AMIT; PACHAURI, KULBHUSHAN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 065216/0649 →
Continuity (2)
Continuation 17714806 · Apr 6, 2022
Related Publication 20240037973A1 · Feb 1, 2024
Cited By (2)
US 12,602,547 US 12,731,424